From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants
Authors
Paper Title
From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants
Publication Info
- Topic area: The impact of mental models on user interaction with AI-based writing assistants.
- Keywords: AI writing assistants, mental models, user behavior, oversight, trust, interaction design, writing quality, human-computer interaction, system safety, user experience.
Background and Problem
- Problem / challenge: While mental models are known to influence user interaction with complex systems, their role in shaping engagement with AI-based writing assistants remains underexplored, particularly in contexts involving system errors.
- Significance: Understanding how mental models affect user oversight and writing outcomes is critical for designing AI systems that support effective and safe user interaction, especially as such systems become ubiquitous in everyday tasks.
- Motivation and related work: Prior research in human-computer interaction and system safety has shown that accurate mental models improve trust, satisfaction, and decision-making. However, gaps remain in understanding how functional versus structural mental models influence user behavior and writing quality in AI-assisted writing.
Solution
- Proposed approach: A controlled experiment comparing the effects of functional and structural mental models on user interaction with an AI-based writing assistant.
- Novelty:
- Empirical analysis of how mental models influence user behavior and writing quality in the presence of AI-generated errors.
- Development of a controlled experimental framework to manipulate and test mental models using instructional priming.
- Insights into the role of interaction affordances in shaping user perceptions of control and ownership.
- Procedure and key techniques:
- Participants (N = 48) were primed with functional or structural mental models via instructional videos.
- Participants completed a cover letter writing task using a modified CoAuthor platform with deliberately injected grammatical errors.
- Behavioral data, final writing quality, and self-reported experiences were analyzed to assess control behavior, writing outcomes, and user perceptions.
Results
- Concrete findings:
- Participants in the structural condition produced letters with more grammatical errors (calibrated corrections: t(46) = −2.316, p < 0.05, d = 0.67).
- Structural participants accepted more erroneous suggestions (errors per word: U = 170.0, p < 0.05, r = 0.352).
- No significant differences were observed in broader writing quality dimensions (relevance, flow, tone/style).
- Structural participants rated the system as easier to use (t(46) = −2.07, p < 0.05, d = 0.60).
- Advantage over baselines:
- Structural mental models increased perceived ease of use but also led to higher error acceptance, suggesting a tradeoff between trust and oversight.
- Interaction affordances (e.g., user control over suggestions) played a more significant role in fostering ownership and control than mental model depth.
- Experiments / evaluation:
- Writing task: Participants wrote a cover letter for a fictional job posting.
- Evaluation metrics: Behavioral logs (e.g., suggestion requests, acceptance ratio), Grammarly correctness scores, qualitative rubric-based assessments, and self-reported surveys.
- Statistical analyses: Independent-samples t-tests, Mann–Whitney U tests, and thematic analysis of interviews.
- Limitations and future work:
- Mental models were assessed at a single point in time and may evolve dynamically.
- Participant variability (e.g., writing proficiency, native language) may have influenced results.
- Future research should explore adaptive scaffolding, progressive disclosure, and interaction techniques to support ongoing mental model calibration.
Summary
This study investigated how functional and structural mental models influence user interaction with AI-based writing assistants. Structural mental models improved perceived ease of use but led to higher acceptance of erroneous suggestions, resulting in more grammatical errors in final outputs. Interaction affordances, rather than mental model depth, were key to fostering user control and ownership. These findings highlight the need for designing AI systems with robust interaction mechanisms and adaptive feedback to support effective oversight and safe use. Future work should explore dynamic mental model calibration and its impact on user behavior in diverse writing contexts.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI Challenges
CHI '23· Human-LLM Collaboration +2
- 71%
Effects of LLM-based Search on Decision Making: Speed, Accuracy, and Overreliance
CHI '25· Human-LLM Collaboration +2
- 71%
Authorship Drift: How Self-Efficacy and Trust Evolve During LLM-Assisted Writing
CHI '26· Human-LLM Collaboration +2
- 71%
Personal Validation Effect in LLMs: Positive AI Responses Bias Perceptions of Validity, Reliability, Personalization, and Usefulness of Fictitious Predictions
CHI '26· Human-LLM Collaboration +2
- 71%
Sensemaking in Multi-Agent LLM Interfaces: How Users Interpret Transparency and Trustworthiness Cues
CHI '26· Human-LLM Collaboration +2
- 71%
Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to Them
CHI '26· Human-LLM Collaboration +2
- 71%
DraftMarks: Enhancing Transparency in Human-AI Co-Writing Through Interactive Skeuomorphic Process Traces
CHI '26· Human-LLM Collaboration +2
- 71%
LLM or Human? Perceptions of Trust and Quality in Research Summaries
CHI '26· Human-LLM Collaboration +2
- 71%
Orality: A Semantic Canvas for Externalizing and Clarifying Thoughts with Speech
CHI '26· Human-LLM Collaboration +2
- 71%
ConstitutionMaker: Interactively Critiquing Large Language Models by Converting Feedback into Principles
IUI '24· Intelligent Voice Assistants (Alexa, Siri, etc.) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)