"Here the GPT made a choice, and every choice can be biased": How Students Critically Engage with LLMs through End-User Auditing Activity

Human-LLM CollaborationAlgorithmic Transparency & AuditabilityUniversity Professors & Researchers

Research Background and Issues

  • What problems or challenges did the authors identify?

    1. The use of large language models (LLMs) is becoming increasingly widespread, but they may generate biased, inaccurate, or unacceptable responses.
    2. Current university policies place the responsibility of verifying the correctness and appropriateness of LLM outputs on students, assuming they possess the necessary AI literacy and skills. This is unfair to students without a technical background.
    3. University students may blindly accept LLM outputs during interactions, lacking critical thinking.
  • Why is this issue important?

    • Content generated by LLMs can have broad and profound impacts on education, learning, and society. Especially when transmitting misinformation or bias, it may put students at risk in terms of academic and ethical judgment.
    • Misguided use of LLMs could exacerbate academic inequality, as students from different backgrounds may have varying abilities to use LLMs effectively.
  • Research Motivation and Related Work

    • Research Motivation: Explore how user auditing can enhance students' critical AI literacy and guide policy and tool design to better serve students.
    • Related Work: Existing studies emphasize the importance of AI literacy in student education, but most classrooms only introduce LLM limitations without providing auditing methods. While there is literature on user-driven algorithm auditing, its application in educational contexts requires further exploration.

Solution

  • What methods or solutions did the authors propose?

    • The authors designed a user auditing tool (PromptAuditor) that combines scenario guidance and specific analysis steps to help students identify and document biases in LLMs.
    • They proposed a system for building user auditing scaffolding, which includes five stages: learning contextual scenarios, hypothesis generation, example demonstration, self-reflection, and comparative learning.
  • What are the innovative aspects of this solution?

    1. Introduced a detailed and structured auditing scaffold to help participants transition from blind exploration to targeted systematic analysis.
    2. Innovatively revealed LLM biases through comparative learning and example discussions, thereby enhancing participants' critical thinking skills.
    3. Conducted empirical research in real educational environments to study how students critically use LLMs, which is rare in existing HCI or AI literacy research.
  • What are the implementation steps and key technologies used?

    • Research Design:
      • Divided into three phases: (1) unguided auditing (based on current university policies); (2) auditing with scaffolding; (3) independent auditing (scaffolding removed).
    • Sample: The study involved 47 university students from various disciplines, conducted in both controlled laboratory environments and natural classroom settings.
    • Technological Applications:
      • Developed and deployed the PromptAuditor web tool, enabling students to experimentally interact with LLMs.
      • Used GPT-3.5 as the underlying model, while the auditing scaffold enhanced students' interaction and analytical capabilities.

Research Findings

  • What specific results were achieved?

    1. Problem Identification: The majority of students found existing university guidance policies on LLMs unclear, difficult to implement, and unfair to non-technical users.
    2. Effectiveness of Scaffolding: With auditing scaffolding, most participants shifted from fragmented exploratory queries to more structured and hypothesis-driven bias audits.
    3. Cognitive Transformation: After using the scaffolding, students not only deepened their exploration of biases within a specific domain but also expanded discussions to include social, historical, and interdisciplinary perspectives.
  • What advantages does this solution have compared to existing ones?

    • Enhanced User Operability: PromptAuditor effectively addresses students' confusion regarding existing policies, teaching them how to systematically identify and document biases.
    • Improved Critical Thinking: Through comparative learning and scenario demonstrations, students transition from being "surface observers" to "in-depth analysts."
    • Adaptability and Fairness: The design is more equitable, considering differences in students' AI knowledge, and is particularly accommodating to non-technical students.
  • What are the experimental or evaluation results?

    1. Surveys revealed that students who had received formal AI literacy education (e.g., completing generative AI courses) performed better in bias detection, but ordinary students also showed significant improvement after receiving scaffold guidance.
    2. Participants were able to articulate and verify the sources of bias more clearly after interacting with GPT, demonstrating effective bias detection in educational contexts.
  • Limitations and Future Directions

    • Limitations:
      1. The study was limited to a single North American university, which may constrain its applicability across different cultural and policy contexts.
      2. Time constraints prevented students from conducting in-depth and sustained audits.
      3. Discussions on sensitive topics (e.g., race, gender) may require additional psychological support.
    • Future Directions:
      1. Develop universal tools for non-technical users applicable across diverse educational environments.
      2. Deepen research on bias perception among students from different social classes or cultural backgrounds.
      3. Explore the possibility of extending AI literacy education to earlier academic stages (e.g., primary and secondary schools) and its impact on intergenerational AI education.

Through such research and improvement measures, the responsible use of LLMs can be further enhanced, promoting a fairer and more comprehensive educational environment.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188978/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713714
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, Algorithmic Transparency & Auditability
work
Professions
University Professors & Researchers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers