Objection Overruled! Lay People can Distinguish Large Language Models from Lawyers, but still Favour Advice from an LLM

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationAI Ethics, Fairness & AccountabilityCybersecurity EngineersAI/ML Researchers & EngineersLawyers & Legal ResearchersHCI Researchers

Research Background and Issues

  • What problems or challenges did the authors identify?
    With the increasing penetration of large language models (LLMs) across various fields, their application in the legal domain has also garnered attention. While LLMs may provide reliable legal advice, the risk of "hallucinations" (i.e., generating inaccurate or false textual content) poses a threat to the public. This study primarily explores:

    1. Are ordinary people more willing to adopt legal advice generated by LLMs compared to advice provided by lawyers?
    2. When the source of the advice is unknown, can ordinary people distinguish between legal advice generated by LLMs and that provided by lawyers?
  • Why is this issue important?
    The use of LLM technology in the legal field could influence public trust in legal advice, and the choices and behaviors of the public in high-risk scenarios could have widespread societal consequences. For example, inaccurate legal advice could lead to poor individual decision-making or negatively impact the judicial system.

  • Research Motivation and Related Work
    Previous research has shown that people exhibit a certain level of trust in algorithm-generated content, although algorithms can lead to errors or misinformation. Surveys have also indicated that the public shows significant interest in using LLM-generated legal advice. However, there has been little in-depth research on whether people are actually willing to adopt such advice and their ability to discern its source.

Proposed Solution

  • What methods or solutions did the authors propose?
    The authors designed three experiments to compare legal advice generated by LLMs and lawyers, investigating ordinary people's behavioral tendencies and ability to discern the source:

    1. Experiment 1: Compare participants' willingness to adopt legal advice when the source is known versus unknown.
    2. Experiment 2: Replicate the "source unknown" condition in Experiment 1 to validate the reliability of the core results.
    3. Experiment 3: Assess participants' accuracy in identifying the source of advice when the source is unknown (using ROC curve analysis).
  • What is innovative about this solution?
    The study focuses on the impact of LLM-generated legal advice in "source unknown" scenarios, revealing a trend in public trust in LLMs compared to traditional professional fields (e.g., lawyers). This design addresses gaps in existing literature regarding public decision-making behavior and discernment ability.

  • What are the implementation steps and key techniques used?

    1. Prepare 18 real legal questions based on UK law, with corresponding legal advice generated by lawyers and ChatGPT-4.
    2. Set experimental conditions, including comparisons between known and unknown sources of legal advice, as well as evaluations of participants' discernment ability.
    3. Use statistical analyses (e.g., ROC curves) and Bayes factor testing to assess participants' willingness and accuracy.
    4. Analyze the reasons behind public preferences for LLM-generated advice and propose solutions to address trust biases.

Research Findings

  • What specific findings were obtained?

    1. Experiments 1 and 2: When the source of legal advice was unknown, participants were significantly more inclined to adopt LLM-generated advice, whereas they showed a stronger preference for lawyer-generated advice when the source was known.
    2. Experiment 3: Although participants demonstrated a higher willingness to adopt LLM-generated advice in "source unknown" scenarios, they were able to significantly distinguish between LLM- and lawyer-generated advice (AUC = 0.59).
  • What advantages does this solution have compared to existing ones?
    This study highlights the coexistence of "high willingness to adopt LLM-generated legal advice and limited discernment ability" among the public, expanding the understanding of AI trust and behavioral tendencies, particularly in high-risk legal scenarios.

  • What were the experimental or evaluation results?

    1. Results from Experiments 1 and 2 showed that LLM-generated advice was more favored in unknown source scenarios (LLM: M=7.32; Lawyer: M=7.02).
    2. ROC analysis in Experiment 3 indicated that participants could distinguish the source of advice, but their accuracy was limited (59% probability, slightly above chance).
  • Limitations and Future Directions

    1. Limitations: Differences in complexity and length between lawyer-generated and LLM-generated content; the LLM sample was limited to ChatGPT-4; the accuracy of the advice was not evaluated.
    2. Future Directions:
      • Explore specific factors influencing public decision-making behavior (e.g., language complexity or social biases).
      • Investigate the potential for collaboration between LLMs and lawyers in generating advice.
      • Design interventions to improve public education and AI content analysis skills to enhance discernment ability.
      • Examine the behavioral effects of different language models (e.g., Google Gemini).

Conclusion

This study demonstrates that ordinary people exhibit an unexpectedly high tendency to adopt LLM-generated legal advice, and although they can distinguish the source of advice, their discernment ability is relatively low. The research emphasizes the potential risks of applying LLMs in high-risk scenarios, such as the "hallucination" problem, and highlights the need to enhance public AI recognition capabilities and relevant policies (e.g., EU AI Act). It also provides future directions for AI research in the legal domain.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189004/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713470
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, AI Ethics, Fairness & Accountability
work
Professions
Cybersecurity Engineers, AI/ML Researchers & Engineers, Lawyers & Legal Researchers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
9 related papers