Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings

Honorable Mention
Human-LLM CollaborationMultilingual & Cross-Cultural Voice InteractionMental Health Apps & Online Support CommunitiesCognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia)Physicians, Nurses & CliniciansCommunity Health WorkersAI/ML Researchers & Engineers

Paper Title

Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings

Publication Info

  • Topic area: Evaluation of Large Language Models (LLMs) in culturally specific healthcare chatbot applications.
  • Keywords: LLM evaluation, healthcare chatbots, community-driven benchmarks, multilingual AI, cultural grounding, civil society, user-centered design, India, AI ethics, inclusive AI.

Background and Problem

  • Problem / challenge: Existing benchmarks for LLMs often rely on generic or translated datasets that fail to capture the cultural, linguistic, and contextual nuances of diverse communities, particularly in high-stakes domains like healthcare.
  • Significance: Ensuring that LLMs are evaluated in ways that reflect the lived realities of underrepresented communities is crucial for their safe and effective adoption in global contexts, especially in healthcare, where inappropriate or insensitive responses can cause harm.
  • Motivation and related work: Prior efforts in LLM evaluation have focused on standardized, often Western-centric benchmarks, which lack ecological validity for non-Western users. While some domain-specific and multilingual benchmarks exist, they often fail to incorporate community-specific cultural contexts. This paper addresses the gap by proposing a community-centered evaluation pipeline.

Solution

  • Proposed approach: Samiksha, a community-driven evaluation pipeline that integrates inputs from civil society organizations (CSOs) and community members to create culturally grounded benchmarks for LLMs.
  • Novelty:
    1. A co-creation process involving CSOs and data workers to design benchmarks that reflect community needs and lived experiences.
    2. A multilingual, culturally contextualized benchmark in three Indian languages (Hindi, Kannada, Malayalam).
    3. A mixed-methods evaluation combining human and LLM-as-judge assessments to analyze model performance and evaluator alignment.
  • Procedure and key techniques:
    1. Phase 1 – Civil Society Consultation: Interviews with five healthcare-focused CSOs to identify key themes, user needs, and evaluation criteria.
    2. Phase 2 – Query Curation: Data workers create and localize healthcare queries based on CSO insights, resulting in 810 original queries and 780 localized queries.
    3. Phase 3 – Response Evaluation: LLM responses are evaluated using rubrics derived from CSO inputs, with both human annotators and LLMs as judges.

Results

  • Concrete findings:
    • Human evaluators provided more nuanced and variable ratings, while LLM judges exhibited compressed, near-ceiling scores with low variance.
    • Across standalone ratings, Qwen3 achieved the highest mean scores, followed by Sarvam-M and Llama-3.1.
    • Comparative evaluations revealed judge-dependent rank shifts, with humans favoring Qwen3 and LLM judges favoring Sarvam-M.
    • Inter-evaluator alignment was weak between humans and LLMs (Pearson r ≈ 0.13) but moderate between LLM judges (r ≈ 0.40).
  • Advantage over baselines: Samiksha outperforms existing benchmarks by incorporating community-specific cultural and linguistic nuances, addressing gaps in ecological validity and inclusivity.
  • Experiments / evaluation:
    • Three multilingual LLMs (Sarvam-M, Qwen3-235B-A22B, Llama-3.1-405B-Instruct) were evaluated on 810 queries in Hindi, Kannada, and Malayalam.
    • Evaluation metrics included clarity, helpfulness, accuracy, and completeness, assessed through standalone and pairwise comparative paradigms.
    • Statistical analyses confirmed significant differences between human and LLM evaluations, highlighting the limitations of automated judges.
  • Limitations and future work:
    • LLM judges struggled with nuanced cultural and linguistic distinctions, necessitating human involvement for critical evaluations.
    • Future work includes scaling the pipeline to additional domains, languages, and geographies, and refining hybrid evaluation workflows that combine automated and human assessments.

Summary

Samiksha introduces a community-driven pipeline for evaluating LLMs in culturally specific healthcare chatbot settings, addressing gaps in inclusivity and ecological validity. By integrating inputs from CSOs and data workers, the approach creates a multilingual benchmark in Hindi, Kannada, and Malayalam, reflecting the lived realities of Indian users. Evaluation results demonstrate the strengths and limitations of human and LLM-as-judge assessments, with significant divergence in their judgments. Samiksha provides a scalable, domain-agnostic template for inclusive LLM evaluation, paving the way for broader adoption across diverse contexts.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222390/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791172
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
8 authors
sell
Subtopics
Human-LLM Collaboration, Multilingual & Cross-Cultural Voice Interaction, Mental Health Apps & Online Support Communities, Cognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia)
work
Professions
Physicians, Nurses & Clinicians, Community Health Workers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers