When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationAI Ethics, Fairness & AccountabilityAI/ML Researchers & EngineersHCI ResearchersPsychiatrists & Psychotherapists

Paper Title

When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being

Publication Info

  • Topic area: Evaluation of AI-generated advice compared to human advice in well-being contexts.
  • Keywords: AI advice, well-being, GPT-4o, GPT-5, human-AI collaboration, Reddit, sycophancy, advice quality, augmentation pipelines, user preferences.

Background and Problem

  • Problem / challenge: The quality of advice generated by large language models (LLMs) for everyday well-being scenarios is unclear, particularly in comparison to human advice from crowdsourced platforms like Reddit.
  • Significance: Understanding the effectiveness of AI-generated advice is critical as LLMs become ubiquitous tools for personal guidance, potentially reshaping how people seek and act on advice.
  • Motivation and related work: Prior research has demonstrated the therapeutic value of empathetic peer communication and crowdsourced advice platforms but highlighted limitations such as delays, low retention, and missed opportunities for sustained support. LLMs offer scalable, personalized advice but are prone to sycophantic behaviors and over-reliance risks. This paper builds on these findings by comparing LLM advice to highly upvoted human advice and exploring human–AI collaboration pipelines.

Solution

  • Proposed approach: Empirical evaluation of advice quality through blinded expert ratings and preference judgments, comparing LLM-generated advice (GPT-4o, GPT-5) to Reddit comments, and testing augmentation pipelines for combining human and AI advice.
  • Novelty:
    1. Blinded, expert-grounded comparison of LLM advice and crowdsourced human advice across six well-being dimensions.
    2. Identification of trade-offs between benchmark gains and practical advice quality, showing GPT-4o outperforming GPT-5.
    3. Exploration of human–AI collaboration pipelines, demonstrating how LLMs can refine human advice and vice versa.
  • Procedure and key techniques:
    • Study-1: Comparison of Reddit comments and LLM-generated advice using expert ratings on six dimensions (e.g., clarity, personalization, sycophancy) and rankings for overall effectiveness and long-term benefit.
    • Study-2: Testing augmentation pipelines where LLMs refine human or AI-originated advice, with or without expert guidance, and evaluating perceived quality and AI-generatedness.
    • Survey: Exploration of user preferences for advice-giving AI personas (Coach vs. Friend) and how preferences vary by trust in AI and baseline well-being.

Results

  • Concrete findings:
    • GPT-4o and GPT-5 outperformed Reddit’s top-rated comments on all six quality dimensions, with GPT-4o ranking higher than GPT-5 in most cases despite GPT-5’s superior benchmarks.
    • Augmenting human advice with LLM edits improved perceived quality, narrowing the gap to LLM-originated advice.
    • Expert-guided LLM edits reduced sycophancy and enhanced perceived humanness but did not consistently outperform LLM-only edits.
  • Advantage over baselines:
    • LLM-generated advice was rated higher than Reddit comments in clarity, effectiveness, personalization, and willingness to seek advice again.
    • Human advice augmented by LLMs was preferred in some scenarios, particularly for overall best rankings.
  • Experiments / evaluation:
    • Study-1: 161 participants evaluated 200 comments across 50 Reddit posts, comparing advice from Reddit’s top-rated comments, GPT-4o, and GPT-5.
    • Study-2: 49 participants evaluated augmented advice from human and LLM seeds across 25 posts, testing LLM-only and LLM+Expert pipelines.
    • Survey: 148 undergraduates assessed qualities and personas of advice-giving AI agents, linking preferences to individual traits like AI trust and well-being.
  • Limitations and future work:
    • Limited generalizability to other domains or clinical contexts.
    • Focus on perceived quality rather than long-term behavioral outcomes.
    • Need for validated frameworks for advice quality and longitudinal studies on adherence and well-being impacts.
    • Exploration of AI disclosure effects and richer content analysis of advice mechanisms.

Summary

This paper provides evidence that frontier LLMs (GPT-4o, GPT-5) outperform crowdsourced human advice on well-being scenarios, with GPT-4o excelling despite GPT-5’s higher benchmarks. Augmentation pipelines show that LLMs can refine human advice effectively, while expert-guided edits reduce sycophancy and enhance humanness. User preferences for advice-giving AI vary by persona and individual traits, suggesting opportunities for persona-sensitive systems. Future work should focus on long-term impacts, safety validation, and hybrid human–AI ecosystems to optimize advice quality and user trust.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222144/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791233
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, AI Ethics, Fairness & Accountability
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, Psychiatrists & Psychotherapists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers