Tell Me What I Missed: Interacting with GPT during Recalling of One-Time Witnessed Events

Human-LLM CollaborationExplainable AI (XAI)Empathy & Emotional DesignHCI ResearchersCognitive ScientistsPsychiatrists & Psychotherapists

Paper Title

Tell Me What I Missed: Interacting with GPT during Recalling of One-Time Witnessed Events

Publication Info

  • Topic area: Investigating the role of GPT in memory recall for eyewitness scenarios.
  • Keywords: GPT, eyewitness memory, recall accuracy, LLM interaction, memory distortion, human-AI collaboration, prompt engineering, subjective interpretation, factual documentation, cognitive processing.

Background and Problem

  • Problem / challenge: While LLMs like GPT are increasingly used for documentation and cognitive support, their impact on memory recall, particularly in eyewitness scenarios, remains underexplored. Concerns include potential distortions in memory and subjective biases introduced by LLM interactions.
  • Significance: Accurate memory recall is critical in high-stakes fields like legal investigations, where distortions can misdirect outcomes. Understanding how LLMs influence memory reconstruction is essential for designing reliable tools.
  • Motivation and related work: Prior research highlights LLMs’ ability to assist in writing and documentation but also warns of risks like false memories and biases. Studies on memory reconstruction emphasize its malleability and susceptibility to external cues, including AI-generated content. This paper builds on these insights to examine how GPT interactions affect memory recall and subjective interpretations in eyewitness contexts.

Solution

  • Proposed approach: A study comparing two GPT interaction conditions—default (natural) and guided (prompted with a standardized eyewitness protocol)—to investigate their effects on memory recall and subjective event interpretation.
  • Novelty:
    1. Demonstrates how GPT interactions influence memory recall and subjective perceptions in eyewitness scenarios.
    2. Highlights differences in user strategies and trust between guided and natural GPT conditions.
    3. Provides empirical insights into the role of prompt design in shaping memory-related outcomes.
    4. Discusses implications for deploying GPT in evidence-based professional fields.
  • Procedure and key techniques:
    • Participants (N=28) watched a 36-second robbery video and used GPT to compose recall statements.
    • Two conditions: natural GPT interaction (default) and guided GPT interaction (prompted with investigative protocols).
    • Post-task assessments included factual recall accuracy, subjective interpretations, and semi-structured interviews.
    • Memory recall was measured using a 50-item factual test, and subjective interpretations were assessed via a 14-item opinion survey.

Results

  • Concrete findings:
    • No significant difference in factual recall accuracy between default (M=24.23, SD=3.19) and guided (M=24.20, SD=5.17) GPT conditions.
    • Participants in the default condition showed greater disfavor toward the intruder (M=8.15, SD=2.70) compared to the guided condition (M=10.00, SD=1.85), indicating heightened affective bias in the default condition.
    • In the default condition, perceived memory clarity was negatively correlated with legitimacy judgments (r=-.68, p=.011) and positively correlated with perceived GPT accuracy (r=.66, p=.014).
    • In the guided condition, perceived clarity aligned with actual recall performance (r=.55, p=.034).
  • Advantage over baselines:
    • Guided GPT interactions aligned subjective clarity with factual recall accuracy, reducing biases compared to the default condition.
    • Prompt engineering mitigated the amplification of affective biases observed in natural GPT interactions.
  • Experiments / evaluation:
    • Participants were randomly assigned to conditions and completed tasks involving video observation, GPT-assisted statement writing, and post-task assessments.
    • Semi-structured interviews revealed diverse interaction strategies, including co-writing, recall coaching, and linguistic refinement.
  • Limitations and future work:
    • Lack of a no-GPT control group to assess the absolute impact of GPT interactions.
    • Short memory consolidation intervals (15 minutes and 1 hour) may not capture long-term memory effects.
    • Future research could explore other LLMs, longer time intervals, and the inclusion of a no-GPT baseline.

Summary

This study investigated how GPT interactions influence memory recall and subjective interpretations in eyewitness scenarios. While factual recall accuracy was similar across default and guided GPT conditions, the default condition amplified affective biases and subjective confidence in GPT’s outputs, even when memory clarity was low. In contrast, guided interactions aligned perceived clarity with actual recall accuracy, reducing biases. Participants developed diverse strategies for using GPT, from co-writing to linguistic refinement. These findings underscore the importance of prompt design in mitigating biases and ensuring reliability in high-stakes applications like legal investigations and journalism. Future work should explore broader contexts and longer-term memory effects.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222251/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791863
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), Empathy & Emotional Design
work
Professions
HCI Researchers, Cognitive Scientists, Psychiatrists & Psychotherapists
article
Content Status
Full text indexed
hub
Related Papers
2 related papers