Beyond Euphemisms: Rethinking LLMs for SRH in Conservative Contexts

Human-LLM CollaborationAI Ethics, Fairness & AccountabilityCognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia)Developing Countries & HCI for Development (HCI4D)Physicians, Nurses & CliniciansPsychiatrists & PsychotherapistsAI/ML Researchers & Engineers

Paper Title

Beyond Euphemisms: Rethinking LLMs for SRH in Conservative Contexts

Publication Info

  • Topic area: Large language models (LLMs) in healthcare communication for sexual and reproductive health (SRH) in conservative, low-resource contexts.
  • Keywords: Large language models, sexual and reproductive health, euphemisms, conservative contexts, healthcare communication, maternal mortality, cultural sensitivity, Roman Urdu, low-resource settings.

Background and Problem

  • Problem / challenge: Current LLMs struggle to interpret indirect, euphemistic, and culturally specific language used in SRH communication, especially in conservative, low-resource contexts like Pakistan. This limits their effectiveness in healthcare interventions.
  • Significance: Addressing communication barriers in SRH can reduce maternal mortality, improve patient engagement, and enhance healthcare delivery in underserved regions.
  • Motivation and related work: Prior work has demonstrated the potential of chatbots and LLMs in healthcare but highlighted persistent issues such as cultural insensitivity, linguistic bias, and limited adaptability to non-English, taboo-laden contexts. This paper builds on these findings by focusing on euphemistic and indirect communication in SRH.

Solution

  • Proposed approach: A two-stage study combining qualitative data collection and LLM evaluation to understand and address indirect SRH communication challenges. The study introduces a Domain-Approach framework for categorizing communication strategies and evaluates five LLMs on their interpretive capabilities.
  • Novelty:
    1. Empirical themes in SRH communication, highlighting indirect language use.
    2. A systematic categorization framework (Domain-Approach) for indirect communication.
    3. Evaluation of LLMs on Roman Urdu SRH prompts, revealing performance gaps.
    4. Design recommendations for culturally situated healthcare interventions using LLMs.
  • Procedure and key techniques:
    • Stage 1: Qualitative data collection through clinician-patient observations, interviews, and focus groups in a charitable hospital in Lahore, Pakistan. Analysis identified euphemisms, myths, and indirect communication patterns.
    • Stage 2: Evaluation of five LLMs (LLaMA 3.2, Gemma 3, GPT-OSS, Claude Sonnet 4, GPT-4o) on 71 Roman Urdu prompts derived from Stage 1. Responses were rated by researchers and a gynecologist for accuracy and cultural appropriateness.

Results

  • Concrete findings:
    • Proprietary models (Claude, GPT-4o) outperformed open-source models (GPT-OSS, Gemma, LLaMA) in interpretive accuracy, with correctness scores of 0.82 and 0.80, respectively.
    • Common failures included misinterpretation of polysemous terms, semantic drift, and overly technical explanations unsuitable for low-literacy patients.
    • Linguistic instability was observed, with models inconsistently switching between Roman Urdu, Hindi, and other scripts.
  • Advantage over baselines:
    • Proprietary models demonstrated higher correctness and coherence compared to open-source systems, but all models struggled with nuanced cultural and linguistic challenges.
  • Experiments / evaluation:
    • 71 Roman Urdu prompts tested across five LLMs, with responses evaluated by researchers and a gynecologist. Metrics included binary correctness, response coherence, and cultural sensitivity.
  • Limitations and future work:
    • Models failed to address polysemy, semantic drift, and gestural communication preferences. Future work should focus on multimodal designs, synchronous terminology management, and active miscommunication protocols.

Summary

This study investigates the challenges of using LLMs for SRH communication in conservative, low-resource contexts, focusing on euphemistic and indirect language patterns. A two-stage methodology revealed systematic communication barriers and evaluated five LLMs on Roman Urdu prompts. Proprietary models performed better than open-source systems but failed to address deeper linguistic complexities like polysemy and semantic drift. The paper proposes design principles for culturally sensitive healthcare interventions, emphasizing synchronous terminology management, default miscommunication protocols, and multimodal designs. These findings contribute to advancing equitable and context-aware AI systems for global health applications.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222689/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791315
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Human-LLM Collaboration, AI Ethics, Fairness & Accountability, Cognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia), Developing Countries & HCI for Development (HCI4D)
work
Professions
Physicians, Nurses & Clinicians, Psychiatrists & Psychotherapists, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers