What Lies Beneath? Exploring the Impact of Underlying AI Model Updates in AI-Infused Systems
Research Background and Issues
- Issues and Challenges: The paper investigates the frequent updates of AI models, particularly those conducted in a "silent" or minimally noticeable manner, leaving users unaware of the changes. These "black-box" model updates can profoundly impact user behavior, task performance, and trust in AI systems, yet this area remains underexplored.
- Significance: As AI models are increasingly applied across diverse domains (e.g., healthcare, finance, image recognition), understanding how updates affect user interaction is crucial to minimizing workflow disruptions and maximizing the benefits of updates.
- Research Motivation and Related Work: While AI updates can introduce new capabilities, they may inadvertently disrupt existing workflows or lead to erroneous outcomes. Although existing guidelines suggest notifying users of modifications, the complexity of comprehensive benchmarking often leads to weakened design and communication practices.
Solution
- Research Methods: The authors conducted two studies to explore the impact of AI model updates on users:
- Online Experiment: Simulated a historical photo recognition task where users were unaware of model updates, observing their ability to detect changes and the effects on behavior and performance.
- Diary Study: Investigated how users perceive different versions of AI models in real-world deployments and explored their preferences regarding model performance.
- Innovations:
- Combined experimental and real-world deployment approaches to explore the practical impacts of AI model updates from multiple dimensions.
- Highlighted users' development of "folk theories," analyzing their unique cognitive interpretations and assumptions about AI behavior.
- Implementation Steps and Key Techniques:
- Simulated a historical photo recognition task in experiments, where users completed eight rounds without knowledge of model differences.
- In real-world platforms, granted users explicit control over two model versions (new and old), allowing them to switch and compare results.
- Collected data on user behavior, accuracy evaluations, model preferences, and task performance metrics.
Research Findings
- Key Results:
- In the experiment, participants' ability to distinguish between model switches was nearly random (accuracy only 48.87%); even though the new model performed better, users failed to fully leverage its capabilities.
- In real-world deployments, most users preferred the new model (due to higher accuracy and quality), but some favored the old model for providing more result options.
- Users developed various conflicting "folk theories" about model behavior (e.g., believing one model focused more on specific facial features or was more sensitive to age).
- Comparison with Existing Solutions:
- Provided a systematic framework combining experimental and real-world studies to demonstrate the complex effects of model updates on user behavior.
- Highlighted trust and efficacy challenges associated with silent updates, differing from the traditional software update domain, which lacks empirical research.
- Experimental Evaluation:
- Experimental Results: The new model significantly improved efficiency (recall increased from 56.05% in the old model to 62.64%), but this did not translate into notable user performance gains.
- Behavior Analysis: Users completed tasks faster in the new model environment, but their speed and scope of interaction adjustments were limited.
- Limitations and Future Directions:
- Limitations: Participants' short-term exposure and familiarity in the experiment may limit ecological validity.
- Future Directions: Suggest developing dynamic communication strategies, such as personalized guidance based on user behavior, and exploring educational interface designs to reduce performance discrepancies in service design.
Conclusion: This study, through two complementary experiments, reveals the significant impact of AI model updates on user behavior and task performance, particularly highlighting that silent updates fail to make improvements perceptible to users. The paper advocates for greater focus on user-centered update communication strategies, such as dynamic behavior-based prompts or multimodal explanation methods, to enable users to better adopt and trust updated AI systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What effects do silent AI model updates have on user behavior and task performance?Category: Mobile Notifications, Timing Management, and Reminder DesignSimilar questionsarrow_forward
- Can users detect AI model updates without notification and effectively use new features?Category: Mobile Notifications, Timing Management, and Reminder DesignSimilar questionsarrow_forward
- How do users form folk theories about AI behavior through cognition, and how do these theories affect interaction?Category: Mobile Notifications, Timing Management, and Reminder DesignSimilar questionsarrow_forward
Practical Problems
1- Users struggle to notice changes after silent AI model updates and cannot effectively use new features.Category: Mobile Notifications, Timing Management, and Reminder DesignSimilar questionsarrow_forward
- 75%
Explanations, Fairness, and Appropriate Reliance in Human-AI Decision-Making
CHI '24· Explainable AI (XAI) +2
- 71%
User Characteristics in Explainable AI: The Rabbit Hole of Personalization?
CHI '24· Explainable AI (XAI) +1
- 71%
I Can Do Better Than Your AI: Expertise and Explanations
IUI '19· Explainable AI (XAI) +1
- 63%
Re-examining Whether, Why, and How Human-AI Interaction Is Uniquely Difficult to Design
CHI '20· Generative AI (Text, Image, Music, Video) +2
- 63%
Knowing About Knowing: An Illusion of Human Competence Can Hinder Appropriate Reliance on AI Systems
CHI '23· Explainable AI (XAI) +2
- 63%
The Metacognitive Demands and Opportunities of Generative AI
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 63%
A Survey of Collaborative Reinforcement Learning: Interactive Methods and Design Patterns
DIS '21· Human-LLM Collaboration +2
- 63%
LESS is More: Rethinking Probabilistic Models of Human Behavior
HRI '20· Brain-Computer Interface (BCI) & Neurofeedback +2
- 63%
Guidance Source Matters: How Guidance from AI, Expert, or a Group of Analysts Impacts Visual Data Preparation and Analysis
IUI '25· Generative AI (Text, Image, Music, Video) +2
- 63%
Optimal Explanations: A Quantitative Model of Human Error in Causal Graph Interpretation
IUI '26· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)