Towards AI-Driven Healthcare: Systematic Optimization, Linguistic Analysis, and Clinicians’ Evaluation of Large Language Models for Smoking Cessation Interventions
Honorable MentionAuthors
Document Title
Towards AI-Driven Healthcare: Systematic Optimization, Linguistic Analysis, and Clinicians’ Evaluation of Large Language Models for Smoking Cessation Interventions
Document Information
- Subject Area: Integration of artificial intelligence and healthcare, smoking cessation interventions
- Keywords: Large Language Models, Information Generation, Computational Linguistic Analysis, Expert Review, Smoking Cessation Interventions
Research Background and Problem
- Main Issues:
- How to optimize Large Language Models (LLMs) to simulate human experts in drafting intervention messages.
- Whether the messages generated by LLMs meet clinical standards for safe application in smoking cessation treatments.
- Significance of the Research:
- Smoking cessation interventions require high-quality message content, which demands significant time and manpower when crafted manually.
- LLM-based generation can accelerate this process, improving intervention efficiency and message quality.
- Motivation and Related Work:
- Existing studies have demonstrated the potential of LLMs in generating information for various health domains, such as COVID-19 vaccine promotion and folic acid-related health information.
- However, research on smoking cessation intervention messages remains underdeveloped, requiring optimization of the generation process and validation of the clinical applicability of the generated content.
Solution
- Method Overview:
The authors systematically explore the process of generating and evaluating smoking cessation intervention messages using LLMs through three studies:
- Prompt Engineering.
- Decoding Optimization.
- Expert Review.
- Innovative Contributions:
- Propose specific strategies for optimizing LLM prompts and decoding methods.
- Introduce a comprehensive evaluation framework combining automated metrics, linguistic attributes, and expert reviews for thorough analysis of information generation.
- Implementation Steps:
- Test five prompt versions across different LLMs, evaluating diversity, quality, and efficiency.
- Optimize eight decoding methods to improve message generation quality, including temperature sampling and top-k sampling.
- Collaborate with smoking cessation treatment experts to review the quality, accuracy, and clinical applicability of LLM-generated messages.
Research Findings
- Specific Outcomes:
- Larger models such as ChatGPT, OPT-13B, and OPT-30B are capable of generating high-quality, credible, and persuasive smoking cessation intervention messages, closely approaching the level of human experts.
- Developed methods for optimizing prompts and decoding, recommending a decoding combination (nucleus sampling combined with top-k sampling and temperature adjustment).
- Comparative Advantages:
- Compared to other LLMs, ChatGPT demonstrated superior message quality, even surpassing human experts.
- Provided an actionable framework for integrating LLM-generated messages into clinical practice.
- Experimental or Evaluation Results:
- Expert reviews indicated that approximately 50% of LLM-generated messages met clinical standards.
- ChatGPT achieved the highest compliance rate (71%), significantly outperforming other models.
- Limitations and Future Directions:
- The sample size for expert reviews was relatively small, potentially introducing representational bias.
- In some cases, the lack of specific contexts affected the effectiveness of message evaluations.
- Future work should integrate more expert review data and provide contextual information to enhance message evaluation accuracy.
Conclusion
This study systematically evaluates LLM applications in high-risk health domains through prompt optimization, decoding optimization, and expert reviews. The findings indicate that LLMs (particularly ChatGPT) can generate high-quality intervention messages, showing promise as a significant auxiliary tool in clinical smoking cessation treatments. However, the study emphasizes that rigorous quality assessments and expert validation remain critical for the safe application of this technology. Future work could expand to model fine-tuning or optimization designs tailored to specific medical scenarios.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can LLMs be optimized to simulate human experts writing smoking cessation intervention messages?Category: Medical Simulation Training and Communication SupportSimilar questionsarrow_forward
- Do LLM-generated smoking cessation intervention messages meet standards for clinically safe application?Category: Medical Simulation Training and Communication SupportSimilar questionsarrow_forward
- Which prompt designs and decoding methods can improve the quality and diversity of LLM-generated messages?Category: Medical Simulation Training and Communication SupportSimilar questionsarrow_forward
Practical Problems
1- Manually writing high-quality smoking cessation intervention messages is time-consuming and labor-intensive.Category: Medical Simulation Training and Communication SupportSimilar questionsarrow_forward
Based on Jaccard similarity of research subtopics & professions (≥60%)