Can LLM-Simulated Practice and Feedback Upskill Human Counselors? A Randomized Study with 90+ Novice Counselors

Honorable Mention
Human-LLM CollaborationMental Health Apps & Online Support CommunitiesMid-Air Haptics (Ultrasonic)Psychiatrists & PsychotherapistsCommunity Health WorkersHCI Researchers

Paper Title

Can LLM-Simulated Practice and Feedback Upskill Human Counselors? A Randomized Study with 90+ Novice Counselors

Publication Info

  • Topic area: Evaluating the impact of LLM-based systems for training novice counselors in client-centered skills.
  • Keywords: LLM, counselor training, empathy, client-centered therapy, AI feedback, skill development, behavioral performance, self-efficacy, therapeutic intentions, randomized study.

Background and Problem

  • Problem / challenge: Training counselors is resource-intensive and difficult to scale. Existing methods rely heavily on human trainers and role-plays, which limit accessibility and scalability. While LLMs offer potential for simulating patients and providing feedback, their effectiveness in improving novice counselor skills remains untested.
  • Significance: Addressing the global shortage of mental health providers requires scalable training solutions. Effective training in foundational counseling skills like empathy and active listening is critical for improving therapy outcomes.
  • Motivation and related work: Prior work has explored LLMs for simulating patient interactions and generating feedback, but evaluations have focused on usability and output quality rather than measurable skill improvement. This study fills the gap by assessing whether LLM-based systems can promote novice skill development.

Solution

  • Proposed approach: CARE, an LLM-based training system combining simulated patient interactions with structured feedback on counseling skills.
  • Novelty:
    1. Integration of realistic LLM-simulated patients with structured, expert-aligned feedback.
    2. Randomized evaluation of skill development in novice counselors (N=94).
    3. Holistic assessment across behavioral performance, self-efficacy, and therapeutic intentions.
    4. Insights into feedback design and its impact on skill adoption.
  • Procedure and key techniques:
    • Participants practiced counseling skills with LLM-simulated patients.
    • Two groups: Practice-only (P) and Practice+Feedback (P+F).
    • Feedback included strengths, areas for improvement, and alternative responses.
    • Skills assessed via fine-tuned LLM classifiers, self-efficacy surveys, and qualitative self-reflections.
    • Randomized pre-post design with three AI patient scenarios.

Results

  • Concrete findings:
    • P+F group improved in reflections (+3.6%, d=0.32, p=0.034) and questions (+6.59%, d=0.36, p=0.018).
    • Empathy declined in the P group (-9.6%, d=-0.52, p<0.001) but remained stable in the P+F group, with a significant between-group difference (d=0.72, p<0.001).
    • Both groups reduced inappropriate suggestions, but feedback was not a significant factor (d=0.02, p=0.910).
  • Advantage over baselines: Feedback prevented empathy decline and promoted client-centered skills, while practice alone led to minimal or negative changes in key areas like empathy.
  • Experiments / evaluation:
    • Participants (N=94) were novice counselors with limited formal training.
    • Measures: Behavioral performance (via LLM classifiers), self-efficacy (CASES-R survey), and qualitative reflections.
    • Study design: Randomized pre-post intervention with two groups (P and P+F).
  • Limitations and future work:
    • Limited to text-based interactions; nonverbal skills not assessed.
    • Short-term study; long-term skill retention and real-world transferability untested.
    • Feedback sometimes perceived as overly critical or repetitive.
    • Future work should explore adaptive patient simulations, multimodal skills, and integration into diverse training contexts.

Summary

This study evaluated CARE, an LLM-based system for training novice counselors, demonstrating that simulated practice combined with structured feedback improves client-centered skills like empathy, reflections, and questions. Practice alone was insufficient and led to empathy decline. Feedback was particularly effective in promoting empathetic listening and discouraging solution-oriented approaches. Participants found the AI feedback constructive and the simulated patients realistic, though some noted challenges with feedback consistency and ecological validity. These findings highlight the potential of LLM-based systems to scale counselor training, though further research is needed to address limitations and explore long-term impacts.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222016/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791821
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, Mental Health Apps & Online Support Communities, Mid-Air Haptics (Ultrasonic)
work
Professions
Psychiatrists & Psychotherapists, Community Health Workers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
5 related papers