A Comparative Analysis of Information Gathering by Chatbots, Questionnaires, and Humans in Clinical Pre-Consultation

Mid-Air Haptics (Ultrasonic)Conversational ChatbotsHuman-LLM CollaborationPhysicians, Nurses & CliniciansPsychiatrists & Psychotherapists

Research Background and Problem

  • Identified Problems or Challenges:
    The authors point out that while large language models (LLMs) have significantly improved the conversational abilities of chatbots, their efficiency in information-gathering tasks remains underexplored. In the context of pre-consultation scenarios (where patients share medical information with intermediaries before seeing a doctor to enhance communication efficiency), the performance of LLM-driven chatbots in information collection compared to static questionnaires and human agents is still unclear. Currently, diagnostic chatbots are prone to decision-making errors due to incomplete input information, but there is a lack of systematic analysis of chatbots' information-gathering capabilities.

  • Significance:
    Accurate information collection is the foundation of diagnosis and decision-making. In clinical scenarios, high-quality information gathering can improve doctor-patient communication, enhance patient engagement, increase medical efficiency, and reduce physicians' documentation burden. Understanding the performance of chatbots in information collection can help improve their design and application.

  • Research Motivation and Related Work:
    The authors propose that while static questionnaires are easy to distribute, they often result in low-quality responses due to questionnaire fatigue. Human interviewers (e.g., nurses) are adaptive but require significant time and human resources. Dynamic chatbots may balance the strengths and weaknesses of both approaches, but current research in this area primarily focuses on chatbot acceptance rather than the quality of their information-gathering capabilities. This motivates the study: how to achieve high-quality information collection with dynamic chatbots.

Solution

  • Methods or Solutions:
    The authors designed a study comparing the efficiency of three pre-consultation information-gathering methods: static questionnaires, GPT-4-based LLM chatbots, and human-operated Wizard-of-Oz scenarios. The study evaluated the quality of information collection based on Grice's conversational maxims (clarity, depth, informativeness, and relevance).

  • Innovative Contributions:

    1. Applied Grice's conversational maxims to evaluate chatbots' information-gathering capabilities, developing novel evaluation metrics (clarity, depth, informativeness, relevance).
    2. Compared the efficiency of information collection by three agents in real clinical pre-consultation scenarios, providing detailed analyses of dynamic follow-up questioning and adaptability.
    3. Proposed specific recommendations for future chatbot design, such as improving dynamic question adjustment and follow-up question design.
  • Implementation Steps and Techniques:

    • Recruited 45 patients who were randomly assigned to one of the three information-gathering agent conditions.
    • Each agent was given 15 standardized questions, but dynamic agents (LLM and Wizard) were allowed to adjust question order, wording, and follow-up questions based on patient responses.
    • Transcripts of the conversations were coded iteratively and evaluated using the four proposed criteria (e.g., clear/unclear).
    • Analyzed the quality of responses to initial questions and the improvement effects of follow-up questions, comparing the overall performance of the three agents.

Research Findings

  • Specific Findings:

    1. Dynamic agents (LLM and Wizard) significantly improved the quality of information collection through dynamic strategies and follow-up questioning, particularly in terms of clarity, depth, informativeness, and relevance of patient responses.
    2. Static questionnaires performed poorly on complex and open-ended questions (e.g., "What medications are you taking?") but were comparable on closed-ended questions (e.g., "What is your pain level?").
    3. Dynamic agents were able to filter redundant information, optimize question order, and appropriately adjust question wording (e.g., expressing empathy or referencing prior information), thereby enhancing the quality of patient responses.
  • Comparison with Existing Solutions and Advantages:

    • Compared to static questionnaires, dynamic agents demonstrated adaptability and flexibility, allowing them to adjust questions based on patient needs for more effective interactions.
    • Compared to chatbots, human-operated Wizards exhibited greater proactivity and follow-up frequency in addressing complex medical issues, but chatbots showed stability and were more suitable for large-scale deployment.
  • Experimental or Evaluation Results:
    The overall satisfactory response rate for dynamic agents (approximately 85%-86%) was significantly higher than that of static questionnaires (79.7%). The frequency of follow-up questions posed by the Wizard was much higher than that of the LLM (16.9% vs. 3.6%). While LLMs could dynamically adjust questions, they still struggled with clarifying ambiguous or multi-topic responses.

  • Limitations and Future Directions:

    • Limitations:
      1. Limited sample size (45 participants), with diverse complexities of medical issues, potentially affecting data consistency.
      2. Did not evaluate doctors' actual satisfaction with the quality of information, relying solely on coding of patient conversations.
      3. The LLM model lacked additional training in specific domains (e.g., medical), which may limit its handling of specialized terminology.
    • Future Directions:
      1. Replicate the study on a larger scale and in standardized medical scenarios.
      2. Develop hybrid models combining chain-of-thought reasoning and retrieval-augmented generation to further enhance multi-topic recording capabilities.
      3. Explore how chatbots can proactively identify and address implicit patient information or potential topics, such as focusing on "critical exclusionary information."

Conclusion and Significance

Research Significance:
This study highlights that dynamic questioning and question adaptability are crucial for enhancing the effectiveness of chatbots in information-gathering tasks. The findings have broad applicability for chatbot design in healthcare and other domains (e.g., customer service, technical support). The research not only provides a framework for developing smarter, more human-centered pre-consultation agents but also offers practical evidence for improving the quality of information collection in automated solutions.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188706/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713613
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Mid-Air Haptics (Ultrasonic), Conversational Chatbots, Human-LLM Collaboration
work
Professions
Physicians, Nurses & Clinicians, Psychiatrists & Psychotherapists
article
Content Status
Full text indexed
hub
Related Papers
2 related papers