CASEbot: A Conversational Agent for Structuring and Personalizing the Design of Self-Experiments in Personal Health
Authors
Paper Title
CASEbot: A Conversational Agent for Structuring and Personalizing the Design of Self-Experiments in Personal Health
Publication Info
- Topic area: Generative AI applications in self-experimentation for personal health.
- Keywords: Self-experimentation, personal health, generative AI, conversational agents, large language models, personalization, safety, experiment design, health informatics, quantified self.
Background and Problem
- Problem / challenge: Existing tools for self-experimentation rely on expert-designed templates for specific conditions, limiting users' ability to create personalized and scientifically rigorous experiments tailored to their own goals.
- Significance: Self-experimentation has the potential to empower individuals to make informed decisions about their health and wellbeing, but current barriers include lack of expertise in experimental design, data collection burden, and safety concerns.
- Motivation and related work: Prior systems like TummyTrials and SleepCoacher addressed barriers by providing pre-designed experiments but lacked scalability and flexibility for personalized designs. Large language models (LLMs) offer potential to bridge these gaps through conversational guidance and adaptive personalization.
Solution
- Proposed approach: CASEbot (Conversation Agent for Self-Experimentation), an LLM-powered chatbot designed to guide users through creating personalized, structured, and safe self-experiments in health domains like physical activity, diet, and sleep.
- Novelty:
- Integration of structure, personalization, and safety principles into a conversational framework for self-experiment design.
- Proactive and reactive personalization through interactive dialogue to adapt experiments to users' goals, habits, and constraints.
- Implementation of safety checks based on reputable sources (e.g., CDC, FDA) to ensure interventions are reasonable and safe.
- Procedure and key techniques:
- Hypothesis formulation: Eliciting user goals and translating them into testable hypotheses.
- Variable identification: Collaboratively defining independent and dependent variables with specificity and feasibility.
- Experiment design: Suggesting appropriate types (e.g., phase-based or randomized) based on temporal relationships between variables.
- Scheduling: Creating tailored experiment calendars based on user routines and constraints.
- Safety assurance: Screening unsafe hypotheses and verifying intervention dosages against established guidelines.
Results
- Concrete findings:
- CASEbot-guided experiments scored higher (average score: 22.31/23) compared to worksheet-guided experiments (average score: 18.98/23), with a 19.20% improvement in quality.
- 92.86% of participants created better experiments using CASEbot.
- Advantage over baselines:
- CASEbot outperformed the worksheet in all structural criteria, particularly in experiment scheduling (+0.929 points), independent variable specificity (+1.000 points), and dependent variable measurability (+0.905 points).
- Participants rated CASEbot higher in ease of use, satisfaction, and learning effectiveness.
- Experiments / evaluation:
- Within-subjects, mixed-methods study with 42 undergraduate participants designing experiments using CASEbot and a static worksheet.
- Evaluation rubric assessed hypotheses, variables, experiment type, and schedule (total score: 23 points).
- Surveys captured subjective experiences; chat transcripts analyzed for qualitative insights.
- Limitations and future work:
- Limited to physical activity, diet, and sleep domains; future work needed to extend to other areas like mental health.
- Participants' high digital literacy may not reflect broader population interaction patterns.
- CASEbot occasionally failed to enforce domain restrictions and provided incorrect information, highlighting the need for stricter guardrails and expert-in-the-loop mechanisms.
Summary
CASEbot leverages generative AI to make self-experimentation more accessible by guiding users through structured, personalized, and safe experiment design. Empirical evaluation demonstrated significant improvements in experiment quality compared to a static worksheet, with participants appreciating its conversational approach and tailored recommendations. However, tensions between guidance and autonomy, as well as challenges in ensuring safety and domain adherence, suggest opportunities for adaptive interactivity and expert validation in future designs. This work underscores the potential of genAI to empower individuals in their health journeys while maintaining scientific rigor and safety.
Research Questions / Practical Problems
Question signals indexed for this paper.
Based on Jaccard similarity of research subtopics & professions (≥60%)