When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
Authors
Paper Title
When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
Publication Info
- Topic area: Evaluation of AI-generated advice compared to human advice in well-being contexts.
- Keywords: AI advice, well-being, GPT-4o, GPT-5, human-AI collaboration, Reddit, sycophancy, advice quality, augmentation pipelines, user preferences.
Background and Problem
- Problem / challenge: The quality of advice generated by large language models (LLMs) for everyday well-being scenarios is unclear, particularly in comparison to human advice from crowdsourced platforms like Reddit.
- Significance: Understanding the effectiveness of AI-generated advice is critical as LLMs become ubiquitous tools for personal guidance, potentially reshaping how people seek and act on advice.
- Motivation and related work: Prior research has demonstrated the therapeutic value of empathetic peer communication and crowdsourced advice platforms but highlighted limitations such as delays, low retention, and missed opportunities for sustained support. LLMs offer scalable, personalized advice but are prone to sycophantic behaviors and over-reliance risks. This paper builds on these findings by comparing LLM advice to highly upvoted human advice and exploring human–AI collaboration pipelines.
Solution
- Proposed approach: Empirical evaluation of advice quality through blinded expert ratings and preference judgments, comparing LLM-generated advice (GPT-4o, GPT-5) to Reddit comments, and testing augmentation pipelines for combining human and AI advice.
- Novelty:
- Blinded, expert-grounded comparison of LLM advice and crowdsourced human advice across six well-being dimensions.
- Identification of trade-offs between benchmark gains and practical advice quality, showing GPT-4o outperforming GPT-5.
- Exploration of human–AI collaboration pipelines, demonstrating how LLMs can refine human advice and vice versa.
- Procedure and key techniques:
- Study-1: Comparison of Reddit comments and LLM-generated advice using expert ratings on six dimensions (e.g., clarity, personalization, sycophancy) and rankings for overall effectiveness and long-term benefit.
- Study-2: Testing augmentation pipelines where LLMs refine human or AI-originated advice, with or without expert guidance, and evaluating perceived quality and AI-generatedness.
- Survey: Exploration of user preferences for advice-giving AI personas (Coach vs. Friend) and how preferences vary by trust in AI and baseline well-being.
Results
- Concrete findings:
- GPT-4o and GPT-5 outperformed Reddit’s top-rated comments on all six quality dimensions, with GPT-4o ranking higher than GPT-5 in most cases despite GPT-5’s superior benchmarks.
- Augmenting human advice with LLM edits improved perceived quality, narrowing the gap to LLM-originated advice.
- Expert-guided LLM edits reduced sycophancy and enhanced perceived humanness but did not consistently outperform LLM-only edits.
- Advantage over baselines:
- LLM-generated advice was rated higher than Reddit comments in clarity, effectiveness, personalization, and willingness to seek advice again.
- Human advice augmented by LLMs was preferred in some scenarios, particularly for overall best rankings.
- Experiments / evaluation:
- Study-1: 161 participants evaluated 200 comments across 50 Reddit posts, comparing advice from Reddit’s top-rated comments, GPT-4o, and GPT-5.
- Study-2: 49 participants evaluated augmented advice from human and LLM seeds across 25 posts, testing LLM-only and LLM+Expert pipelines.
- Survey: 148 undergraduates assessed qualities and personas of advice-giving AI agents, linking preferences to individual traits like AI trust and well-being.
- Limitations and future work:
- Limited generalizability to other domains or clinical contexts.
- Focus on perceived quality rather than long-term behavioral outcomes.
- Need for validated frameworks for advice quality and longitudinal studies on adherence and well-being impacts.
- Exploration of AI disclosure effects and richer content analysis of advice mechanisms.
Summary
This paper provides evidence that frontier LLMs (GPT-4o, GPT-5) outperform crowdsourced human advice on well-being scenarios, with GPT-4o excelling despite GPT-5’s higher benchmarks. Augmentation pipelines show that LLMs can refine human advice effectively, while expert-guided edits reduce sycophancy and enhance humanness. User preferences for advice-giving AI vary by persona and individual traits, suggesting opportunities for persona-sensitive systems. Future work should focus on long-term impacts, safety validation, and hybrid human–AI ecosystems to optimize advice quality and user trust.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Interaction Context Often Increases Sycophancy in LLMs
CHI '26· Human-LLM Collaboration +2
- 71%
Are Two Heads Better Than One in AI-Assisted Decision Making? Comparing the Behavior and Performance of Groups and Individuals in Human-AI Collaborative Recidivism Risk Assessment
CHI '23· Human-LLM Collaboration +2
- 71%
What is Human-Centered about Human-Centered AI? A Map of the Research Landscape
CHI '23· Human-LLM Collaboration +2
- 71%
Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions
CHI '24· Human-LLM Collaboration +2
- 71%
Understanding Socio-technical Factors Configuring AI Non-Use in UX Work Practices
CHI '25· Human-LLM Collaboration +2
- 71%
Framing Responsible Design of AI for Mental Well-Being: AI as Primary Care, Nutritional Supplement, or Yoga Instructor?
CHI '26· Human-LLM Collaboration +2
- 71%
FAIR: Framing AI’s Role in Programming Competitions — Understanding How LLMs Are Changing the Game in Competitive Programming
CHI '26· Human-LLM Collaboration +2
- 71%
Understanding Compliance and Conversion Dynamics in Multi-Agent Collectives
CHI '26· Human-LLM Collaboration +2
- 71%
Large Language Models in Peer-Run Community Behavioral Health Services: Understanding Peer Specialists and Service Users’ Perspectives on Opportunities, Risks, and Mitigation Strategies
CHI '26· Human-LLM Collaboration +2
- 71%
Who Does What? Archetypes of Roles Assigned to LLMs During Human-AI Decision-Making
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)