Towards AI-Driven Healthcare: Systematic Optimization, Linguistic Analysis, and Clinicians’ Evaluation of Large Language Models for Smoking Cessation Interventions

Honorable Mention
Mid-Air Haptics (Ultrasonic)Human-LLM CollaborationAI-Assisted Decision-Making & AutomationPhysicians, Nurses & CliniciansAI/ML Researchers & Engineers

Document Title

Towards AI-Driven Healthcare: Systematic Optimization, Linguistic Analysis, and Clinicians’ Evaluation of Large Language Models for Smoking Cessation Interventions

Document Information

  • Subject Area: Integration of artificial intelligence and healthcare, smoking cessation interventions
  • Keywords: Large Language Models, Information Generation, Computational Linguistic Analysis, Expert Review, Smoking Cessation Interventions

Research Background and Problem

  • Main Issues:
    1. How to optimize Large Language Models (LLMs) to simulate human experts in drafting intervention messages.
    2. Whether the messages generated by LLMs meet clinical standards for safe application in smoking cessation treatments.
  • Significance of the Research:
    • Smoking cessation interventions require high-quality message content, which demands significant time and manpower when crafted manually.
    • LLM-based generation can accelerate this process, improving intervention efficiency and message quality.
  • Motivation and Related Work:
    • Existing studies have demonstrated the potential of LLMs in generating information for various health domains, such as COVID-19 vaccine promotion and folic acid-related health information.
    • However, research on smoking cessation intervention messages remains underdeveloped, requiring optimization of the generation process and validation of the clinical applicability of the generated content.

Solution

  • Method Overview: The authors systematically explore the process of generating and evaluating smoking cessation intervention messages using LLMs through three studies:
    1. Prompt Engineering.
    2. Decoding Optimization.
    3. Expert Review.
  • Innovative Contributions:
    • Propose specific strategies for optimizing LLM prompts and decoding methods.
    • Introduce a comprehensive evaluation framework combining automated metrics, linguistic attributes, and expert reviews for thorough analysis of information generation.
  • Implementation Steps:
    1. Test five prompt versions across different LLMs, evaluating diversity, quality, and efficiency.
    2. Optimize eight decoding methods to improve message generation quality, including temperature sampling and top-k sampling.
    3. Collaborate with smoking cessation treatment experts to review the quality, accuracy, and clinical applicability of LLM-generated messages.

Research Findings

  • Specific Outcomes:
    • Larger models such as ChatGPT, OPT-13B, and OPT-30B are capable of generating high-quality, credible, and persuasive smoking cessation intervention messages, closely approaching the level of human experts.
    • Developed methods for optimizing prompts and decoding, recommending a decoding combination (nucleus sampling combined with top-k sampling and temperature adjustment).
  • Comparative Advantages:
    • Compared to other LLMs, ChatGPT demonstrated superior message quality, even surpassing human experts.
    • Provided an actionable framework for integrating LLM-generated messages into clinical practice.
  • Experimental or Evaluation Results:
    • Expert reviews indicated that approximately 50% of LLM-generated messages met clinical standards.
    • ChatGPT achieved the highest compliance rate (71%), significantly outperforming other models.
  • Limitations and Future Directions:
    • The sample size for expert reviews was relatively small, potentially introducing representational bias.
    • In some cases, the lack of specific contexts affected the effectiveness of message evaluations.
    • Future work should integrate more expert review data and provide contextual information to enhance message evaluation accuracy.

Conclusion

This study systematically evaluates LLM applications in high-risk health domains through prompt optimization, decoding optimization, and expert reviews. The findings indicate that LLMs (particularly ChatGPT) can generate high-quality intervention messages, showing promise as a significant auxiliary tool in clinical smoking cessation treatments. However, the study emphasizes that rigorous quality assessments and expert validation remain critical for the safe application of this technology. Future work could expand to model fine-tuning or optimization designs tailored to specific medical scenarios.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147450/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3641965
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
Honorable Mention
group
Authors
8 authors
sell
Subtopics
Mid-Air Haptics (Ultrasonic), Human-LLM Collaboration, AI-Assisted Decision-Making & Automation
work
Professions
Physicians, Nurses & Clinicians, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers