CASEbot: A Conversational Agent for Structuring and Personalizing the Design of Self-Experiments in Personal Health

Generative AI (Text, Image, Music, Video)Human-LLM CollaborationBehavior Change & Reflection TechnologyHealth Self-TrackingPhysicians, Nurses & CliniciansPsychiatrists & PsychotherapistsAI/ML Researchers & Engineers

Paper Title

CASEbot: A Conversational Agent for Structuring and Personalizing the Design of Self-Experiments in Personal Health

Publication Info

  • Topic area: Generative AI applications in self-experimentation for personal health.
  • Keywords: Self-experimentation, personal health, generative AI, conversational agents, large language models, personalization, safety, experiment design, health informatics, quantified self.

Background and Problem

  • Problem / challenge: Existing tools for self-experimentation rely on expert-designed templates for specific conditions, limiting users' ability to create personalized and scientifically rigorous experiments tailored to their own goals.
  • Significance: Self-experimentation has the potential to empower individuals to make informed decisions about their health and wellbeing, but current barriers include lack of expertise in experimental design, data collection burden, and safety concerns.
  • Motivation and related work: Prior systems like TummyTrials and SleepCoacher addressed barriers by providing pre-designed experiments but lacked scalability and flexibility for personalized designs. Large language models (LLMs) offer potential to bridge these gaps through conversational guidance and adaptive personalization.

Solution

  • Proposed approach: CASEbot (Conversation Agent for Self-Experimentation), an LLM-powered chatbot designed to guide users through creating personalized, structured, and safe self-experiments in health domains like physical activity, diet, and sleep.
  • Novelty:
    1. Integration of structure, personalization, and safety principles into a conversational framework for self-experiment design.
    2. Proactive and reactive personalization through interactive dialogue to adapt experiments to users' goals, habits, and constraints.
    3. Implementation of safety checks based on reputable sources (e.g., CDC, FDA) to ensure interventions are reasonable and safe.
  • Procedure and key techniques:
    • Hypothesis formulation: Eliciting user goals and translating them into testable hypotheses.
    • Variable identification: Collaboratively defining independent and dependent variables with specificity and feasibility.
    • Experiment design: Suggesting appropriate types (e.g., phase-based or randomized) based on temporal relationships between variables.
    • Scheduling: Creating tailored experiment calendars based on user routines and constraints.
    • Safety assurance: Screening unsafe hypotheses and verifying intervention dosages against established guidelines.

Results

  • Concrete findings:
    • CASEbot-guided experiments scored higher (average score: 22.31/23) compared to worksheet-guided experiments (average score: 18.98/23), with a 19.20% improvement in quality.
    • 92.86% of participants created better experiments using CASEbot.
  • Advantage over baselines:
    • CASEbot outperformed the worksheet in all structural criteria, particularly in experiment scheduling (+0.929 points), independent variable specificity (+1.000 points), and dependent variable measurability (+0.905 points).
    • Participants rated CASEbot higher in ease of use, satisfaction, and learning effectiveness.
  • Experiments / evaluation:
    • Within-subjects, mixed-methods study with 42 undergraduate participants designing experiments using CASEbot and a static worksheet.
    • Evaluation rubric assessed hypotheses, variables, experiment type, and schedule (total score: 23 points).
    • Surveys captured subjective experiences; chat transcripts analyzed for qualitative insights.
  • Limitations and future work:
    • Limited to physical activity, diet, and sleep domains; future work needed to extend to other areas like mental health.
    • Participants' high digital literacy may not reflect broader population interaction patterns.
    • CASEbot occasionally failed to enforce domain restrictions and provided incorrect information, highlighting the need for stricter guardrails and expert-in-the-loop mechanisms.

Summary

CASEbot leverages generative AI to make self-experimentation more accessible by guiding users through structured, personalized, and safe experiment design. Empirical evaluation demonstrated significant improvements in experiment quality compared to a static worksheet, with participants appreciating its conversational approach and tailored recommendations. However, tensions between guidance and autonomy, as well as challenges in ensuring safety and domain adherence, suggest opportunities for adaptive interactivity and expert validation in future designs. This work underscores the potential of genAI to empower individuals in their health journeys while maintaining scientific rigor and safety.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223295/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791551
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration, Behavior Change & Reflection Technology, Health Self-Tracking
work
Professions
Physicians, Nurses & Clinicians, Psychiatrists & Psychotherapists, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers