Tell Me What I Missed: Interacting with GPT during Recalling of One-Time Witnessed Events
Authors
Paper Title
Tell Me What I Missed: Interacting with GPT during Recalling of One-Time Witnessed Events
Publication Info
- Topic area: Investigating the role of GPT in memory recall for eyewitness scenarios.
- Keywords: GPT, eyewitness memory, recall accuracy, LLM interaction, memory distortion, human-AI collaboration, prompt engineering, subjective interpretation, factual documentation, cognitive processing.
Background and Problem
- Problem / challenge: While LLMs like GPT are increasingly used for documentation and cognitive support, their impact on memory recall, particularly in eyewitness scenarios, remains underexplored. Concerns include potential distortions in memory and subjective biases introduced by LLM interactions.
- Significance: Accurate memory recall is critical in high-stakes fields like legal investigations, where distortions can misdirect outcomes. Understanding how LLMs influence memory reconstruction is essential for designing reliable tools.
- Motivation and related work: Prior research highlights LLMs’ ability to assist in writing and documentation but also warns of risks like false memories and biases. Studies on memory reconstruction emphasize its malleability and susceptibility to external cues, including AI-generated content. This paper builds on these insights to examine how GPT interactions affect memory recall and subjective interpretations in eyewitness contexts.
Solution
- Proposed approach: A study comparing two GPT interaction conditions—default (natural) and guided (prompted with a standardized eyewitness protocol)—to investigate their effects on memory recall and subjective event interpretation.
- Novelty:
- Demonstrates how GPT interactions influence memory recall and subjective perceptions in eyewitness scenarios.
- Highlights differences in user strategies and trust between guided and natural GPT conditions.
- Provides empirical insights into the role of prompt design in shaping memory-related outcomes.
- Discusses implications for deploying GPT in evidence-based professional fields.
- Procedure and key techniques:
- Participants (N=28) watched a 36-second robbery video and used GPT to compose recall statements.
- Two conditions: natural GPT interaction (default) and guided GPT interaction (prompted with investigative protocols).
- Post-task assessments included factual recall accuracy, subjective interpretations, and semi-structured interviews.
- Memory recall was measured using a 50-item factual test, and subjective interpretations were assessed via a 14-item opinion survey.
Results
- Concrete findings:
- No significant difference in factual recall accuracy between default (M=24.23, SD=3.19) and guided (M=24.20, SD=5.17) GPT conditions.
- Participants in the default condition showed greater disfavor toward the intruder (M=8.15, SD=2.70) compared to the guided condition (M=10.00, SD=1.85), indicating heightened affective bias in the default condition.
- In the default condition, perceived memory clarity was negatively correlated with legitimacy judgments (r=-.68, p=.011) and positively correlated with perceived GPT accuracy (r=.66, p=.014).
- In the guided condition, perceived clarity aligned with actual recall performance (r=.55, p=.034).
- Advantage over baselines:
- Guided GPT interactions aligned subjective clarity with factual recall accuracy, reducing biases compared to the default condition.
- Prompt engineering mitigated the amplification of affective biases observed in natural GPT interactions.
- Experiments / evaluation:
- Participants were randomly assigned to conditions and completed tasks involving video observation, GPT-assisted statement writing, and post-task assessments.
- Semi-structured interviews revealed diverse interaction strategies, including co-writing, recall coaching, and linguistic refinement.
- Limitations and future work:
- Lack of a no-GPT control group to assess the absolute impact of GPT interactions.
- Short memory consolidation intervals (15 minutes and 1 hour) may not capture long-term memory effects.
- Future research could explore other LLMs, longer time intervals, and the inclusion of a no-GPT baseline.
Summary
This study investigated how GPT interactions influence memory recall and subjective interpretations in eyewitness scenarios. While factual recall accuracy was similar across default and guided GPT conditions, the default condition amplified affective biases and subjective confidence in GPT’s outputs, even when memory clarity was low. In contrast, guided interactions aligned perceived clarity with actual recall accuracy, reducing biases. Participants developed diverse strategies for using GPT, from co-writing to linguistic refinement. These findings underscore the importance of prompt design in mitigating biases and ensuring reliability in high-stakes applications like legal investigations and journalism. Future work should explore broader contexts and longer-term memory effects.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
EvAlignUX: Advancing UX Evaluation through LLM-Supported Metrics Exploration
CHI '25· Human-LLM Collaboration +1
- 63%
Semantic See-through Goggles: Wearing Linguistic Virtual Reality in (Artificial Intelligence)
IUI '26· Generative AI (Text, Image, Music, Video) +3
Based on Jaccard similarity of research subtopics & professions (≥60%)