What Lies Beneath? Exploring the Impact of Underlying AI Model Updates in AI-Infused Systems

Generative AI (Text, Image, Music, Video)Explainable AI (XAI)AI-Assisted Decision-Making & AutomationAI/ML Researchers & EngineersHCI ResearchersCognitive ScientistsSociologists & Anthropologists

Research Background and Issues

  • Issues and Challenges: The paper investigates the frequent updates of AI models, particularly those conducted in a "silent" or minimally noticeable manner, leaving users unaware of the changes. These "black-box" model updates can profoundly impact user behavior, task performance, and trust in AI systems, yet this area remains underexplored.
  • Significance: As AI models are increasingly applied across diverse domains (e.g., healthcare, finance, image recognition), understanding how updates affect user interaction is crucial to minimizing workflow disruptions and maximizing the benefits of updates.
  • Research Motivation and Related Work: While AI updates can introduce new capabilities, they may inadvertently disrupt existing workflows or lead to erroneous outcomes. Although existing guidelines suggest notifying users of modifications, the complexity of comprehensive benchmarking often leads to weakened design and communication practices.

Solution

  • Research Methods: The authors conducted two studies to explore the impact of AI model updates on users:
    1. Online Experiment: Simulated a historical photo recognition task where users were unaware of model updates, observing their ability to detect changes and the effects on behavior and performance.
    2. Diary Study: Investigated how users perceive different versions of AI models in real-world deployments and explored their preferences regarding model performance.
  • Innovations:
    • Combined experimental and real-world deployment approaches to explore the practical impacts of AI model updates from multiple dimensions.
    • Highlighted users' development of "folk theories," analyzing their unique cognitive interpretations and assumptions about AI behavior.
  • Implementation Steps and Key Techniques:
    1. Simulated a historical photo recognition task in experiments, where users completed eight rounds without knowledge of model differences.
    2. In real-world platforms, granted users explicit control over two model versions (new and old), allowing them to switch and compare results.
    3. Collected data on user behavior, accuracy evaluations, model preferences, and task performance metrics.

Research Findings

  • Key Results:
    1. In the experiment, participants' ability to distinguish between model switches was nearly random (accuracy only 48.87%); even though the new model performed better, users failed to fully leverage its capabilities.
    2. In real-world deployments, most users preferred the new model (due to higher accuracy and quality), but some favored the old model for providing more result options.
    3. Users developed various conflicting "folk theories" about model behavior (e.g., believing one model focused more on specific facial features or was more sensitive to age).
  • Comparison with Existing Solutions:
    • Provided a systematic framework combining experimental and real-world studies to demonstrate the complex effects of model updates on user behavior.
    • Highlighted trust and efficacy challenges associated with silent updates, differing from the traditional software update domain, which lacks empirical research.
  • Experimental Evaluation:
    • Experimental Results: The new model significantly improved efficiency (recall increased from 56.05% in the old model to 62.64%), but this did not translate into notable user performance gains.
    • Behavior Analysis: Users completed tasks faster in the new model environment, but their speed and scope of interaction adjustments were limited.
  • Limitations and Future Directions:
    1. Limitations: Participants' short-term exposure and familiarity in the experiment may limit ecological validity.
    2. Future Directions: Suggest developing dynamic communication strategies, such as personalized guidance based on user behavior, and exploring educational interface designs to reduce performance discrepancies in service design.

Conclusion: This study, through two complementary experiments, reveals the significant impact of AI model updates on user behavior and task performance, particularly highlighting that silent updates fail to make improvements perceptible to users. The paper advocates for greater focus on user-centered update communication strategies, such as dynamic behavior-based prompts or multimodal explanation methods, to enable users to better adopt and trust updated AI systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188738/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713751
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Explainable AI (XAI), AI-Assisted Decision-Making & Automation
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, Cognitive Scientists, Sociologists & Anthropologists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers