Objection Overruled! Lay People can Distinguish Large Language Models from Lawyers, but still Favour Advice from an LLM
Authors
Research Background and Issues
-
What problems or challenges did the authors identify?
With the increasing penetration of large language models (LLMs) across various fields, their application in the legal domain has also garnered attention. While LLMs may provide reliable legal advice, the risk of "hallucinations" (i.e., generating inaccurate or false textual content) poses a threat to the public. This study primarily explores:- Are ordinary people more willing to adopt legal advice generated by LLMs compared to advice provided by lawyers?
- When the source of the advice is unknown, can ordinary people distinguish between legal advice generated by LLMs and that provided by lawyers?
-
Why is this issue important?
The use of LLM technology in the legal field could influence public trust in legal advice, and the choices and behaviors of the public in high-risk scenarios could have widespread societal consequences. For example, inaccurate legal advice could lead to poor individual decision-making or negatively impact the judicial system. -
Research Motivation and Related Work
Previous research has shown that people exhibit a certain level of trust in algorithm-generated content, although algorithms can lead to errors or misinformation. Surveys have also indicated that the public shows significant interest in using LLM-generated legal advice. However, there has been little in-depth research on whether people are actually willing to adopt such advice and their ability to discern its source.
Proposed Solution
-
What methods or solutions did the authors propose?
The authors designed three experiments to compare legal advice generated by LLMs and lawyers, investigating ordinary people's behavioral tendencies and ability to discern the source:- Experiment 1: Compare participants' willingness to adopt legal advice when the source is known versus unknown.
- Experiment 2: Replicate the "source unknown" condition in Experiment 1 to validate the reliability of the core results.
- Experiment 3: Assess participants' accuracy in identifying the source of advice when the source is unknown (using ROC curve analysis).
-
What is innovative about this solution?
The study focuses on the impact of LLM-generated legal advice in "source unknown" scenarios, revealing a trend in public trust in LLMs compared to traditional professional fields (e.g., lawyers). This design addresses gaps in existing literature regarding public decision-making behavior and discernment ability. -
What are the implementation steps and key techniques used?
- Prepare 18 real legal questions based on UK law, with corresponding legal advice generated by lawyers and ChatGPT-4.
- Set experimental conditions, including comparisons between known and unknown sources of legal advice, as well as evaluations of participants' discernment ability.
- Use statistical analyses (e.g., ROC curves) and Bayes factor testing to assess participants' willingness and accuracy.
- Analyze the reasons behind public preferences for LLM-generated advice and propose solutions to address trust biases.
Research Findings
-
What specific findings were obtained?
- Experiments 1 and 2: When the source of legal advice was unknown, participants were significantly more inclined to adopt LLM-generated advice, whereas they showed a stronger preference for lawyer-generated advice when the source was known.
- Experiment 3: Although participants demonstrated a higher willingness to adopt LLM-generated advice in "source unknown" scenarios, they were able to significantly distinguish between LLM- and lawyer-generated advice (AUC = 0.59).
-
What advantages does this solution have compared to existing ones?
This study highlights the coexistence of "high willingness to adopt LLM-generated legal advice and limited discernment ability" among the public, expanding the understanding of AI trust and behavioral tendencies, particularly in high-risk legal scenarios. -
What were the experimental or evaluation results?
- Results from Experiments 1 and 2 showed that LLM-generated advice was more favored in unknown source scenarios (LLM: M=7.32; Lawyer: M=7.02).
- ROC analysis in Experiment 3 indicated that participants could distinguish the source of advice, but their accuracy was limited (59% probability, slightly above chance).
-
Limitations and Future Directions
- Limitations: Differences in complexity and length between lawyer-generated and LLM-generated content; the LLM sample was limited to ChatGPT-4; the accuracy of the advice was not evaluated.
- Future Directions:
- Explore specific factors influencing public decision-making behavior (e.g., language complexity or social biases).
- Investigate the potential for collaboration between LLMs and lawyers in generating advice.
- Design interventions to improve public education and AI content analysis skills to enhance discernment ability.
- Examine the behavioral effects of different language models (e.g., Google Gemini).
Conclusion
This study demonstrates that ordinary people exhibit an unexpectedly high tendency to adopt LLM-generated legal advice, and although they can distinguish the source of advice, their discernment ability is relatively low. The research emphasizes the potential risks of applying LLMs in high-risk scenarios, such as the "hallucination" problem, and highlights the need to enhance public AI recognition capabilities and relevant policies (e.g., EU AI Act). It also provides future directions for AI research in the legal domain.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
2- Are laypeople more willing to adopt legal advice generated by large language models (LLMs) than advice provided by lawyers?Category: Medical Risk Explanation and Hypothesis ExplorationSimilar questionsarrow_forward
- When the source of legal advice is unknown, can laypeople distinguish LLM-generated advice from lawyer-provided advice?Category: Medical Risk Explanation and Hypothesis ExplorationSimilar questionsarrow_forward
Practical Problems
1- Laypeople may trust inaccurate LLM-generated legal advice in high-risk scenarios.Category: Medical Risk Explanation and Hypothesis ExplorationSimilar questionsarrow_forward
- 71%
Interaction Context Often Increases Sycophancy in LLMs
CHI '26· Human-LLM Collaboration +2
- 63%
Are Two Heads Better Than One in AI-Assisted Decision Making? Comparing the Behavior and Performance of Groups and Individuals in Human-AI Collaborative Recidivism Risk Assessment
CHI '23· Human-LLM Collaboration +2
- 63%
What is Human-Centered about Human-Centered AI? A Map of the Research Landscape
CHI '23· Human-LLM Collaboration +2
- 63%
Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions
CHI '24· Human-LLM Collaboration +2
- 63%
A Matter of Perspective(s): Contrasting Human and LLM Argumentation in Subjective Decision-Making on Subtle Sexism
CHI '25· Human-LLM Collaboration +2
- 63%
Understanding Socio-technical Factors Configuring AI Non-Use in UX Work Practices
CHI '25· Human-LLM Collaboration +2
- 63%
When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
CHI '26· Human-LLM Collaboration +2
- 63%
FAIR: Framing AI’s Role in Programming Competitions — Understanding How LLMs Are Changing the Game in Competitive Programming
CHI '26· Human-LLM Collaboration +2
- 63%
Understanding Compliance and Conversion Dynamics in Multi-Agent Collectives
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)