Quantifying the Novelty Bias when Evaluating Interactive Prototypes
Honorable MentionAuthors
Paper Title
Quantifying the Novelty Bias when Evaluating Interactive Prototypes
Publication Info
- Topic area: Human-computer interaction (HCI) evaluation and novelty bias.
- Keywords: Novelty effect, human-computer interaction, subjective ratings, performance metrics, technology readiness, prototype evaluation, placebo effect, user studies, bias quantification, interaction design.
Background and Problem
- Problem / challenge: HCI evaluations often rely on user judgments to assess new technologies, but novelty bias—favoring what is labeled as “new”—can distort these judgments, leading to inflated preferences and ratings that may not align with actual performance.
- Significance: Understanding and quantifying novelty bias is crucial for ensuring the validity of HCI evaluations, particularly as they inform design decisions and claims of improvement.
- Motivation and related work: Prior research has acknowledged novelty effects in HCI but has not quantified their impact on subjective ratings and objective performance. Related studies on placebo effects and evaluation validity highlight the need to isolate novelty framing as a distinct biasing force.
Solution
- Proposed approach: A controlled experimental study to quantify the impact of novelty labeling on subjective preference and objective performance across four technology types: mice, keyboards, search engines, and AI chatbots.
- Novelty:
- Demonstrates causally that labeling a functionally identical system as “new” shifts preferences and subjective ratings.
- Shows that these perceptual shifts can outweigh or misalign with objective performance metrics.
- Clarifies the role of technology readiness (TR) in moderating baseline performance but not shielding judgments from novelty bias.
- Procedure and key techniques:
- Participants interacted with functionally identical technology pairs differing only in cosmetic features and “old”/“new” labels.
- Tasks included pointing, typing, searching, and chatting, with measures of performance (e.g., throughput, error rates) and subjective ratings (e.g., Likert scales).
- Extensive counterbalancing ensured no systematic bias from order effects or cosmetic preferences.
- Statistical models analyzed trial-level and participant-level outcomes, including interactions with TR.
Results
- Concrete findings:
- Up to 77% of participants preferred the “new” versions, with subjective ratings inflated by up to 11.6% for preferred AI chatbots and 7.1% for the “new” search engine.
- Performance differences were modest: 9.7% fewer misses with the “new” mouse and up to 7.2% lower error rates with the “new” keyboard.
- Throughput, movement time, browsing time, and number of messages exchanged showed little change.
- Advantage over baselines: Novelty labeling consistently shifted preferences and ratings, but objective performance metrics showed limited improvement, highlighting a misalignment between perceived and actual benefits.
- Experiments / evaluation:
- Sample size: 48 participants across diverse backgrounds.
- Technologies tested: mice, keyboards, search engines, AI chatbots.
- Metrics: speed, accuracy, subjective ratings, and TR scores.
- Statistical methods: mixed-effects models, nonparametric tests, and binomial tests for preference.
- Limitations and future work:
- Results focus on first-use, structured lab settings; longer-term and field studies are needed to assess novelty effects over time.
- Manipulation relied on explicit labels; future work could explore subtler novelty cues.
- Broader technology classes and higher-level tasks (e.g., collaboration, creative work) may reveal different sensitivities to novelty.
Summary
This study quantifies the impact of novelty labeling on subjective preference and objective performance across four technology types. Results show that novelty framing strongly influences preference and subjective ratings, with limited effects on performance metrics such as speed and accuracy. Technology readiness predicts baseline skill but does not mitigate novelty bias. These findings highlight the methodological challenge of separating genuine improvements from novelty-driven perceptions in HCI evaluations. Future research should explore longitudinal effects, subtler novelty cues, and broader technology domains to refine evaluation practices and ensure reliable conclusions about system quality.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Examining Design Choices of Questionnaires in VR User Studies
CHI '20· Immersion & Presence Research +2
- 67%
How Do We Measure That?! Quick Scale Development
CHI '18· User Research Methods (Interviews, Surveys, Observation) +1
- 67%
A Bermuda Triangle? - A Review of Method Application and Triangulation in User Experience Evaluation
CHI '18· User Research Methods (Interviews, Surveys, Observation) +1
- 67%
Designing with the Mind in Mind: The Psychological Basis for UI Design Guidelines
CHI '18· User Research Methods (Interviews, Surveys, Observation) +1
- 67%
Balanced Interaction Design
CHI '18· User Research Methods (Interviews, Surveys, Observation) +1
- 67%
The Meaning of Interactivity—Some Proposals for Definitions and Measures
CHI '18· User Research Methods (Interviews, Surveys, Observation) +1
- 67%
Technology Acceptance and User Experience: A Review of the Experiential Component in HCI
CHI '18· User Research Methods (Interviews, Surveys, Observation) +1
- 67%
Ethical Mediation in UX Practice
CHI '19· User Research Methods (Interviews, Surveys, Observation) +1
- 67%
A Practice-Led Account of the Conceptual Evolution of UX Knowledge
CHI '19· User Research Methods (Interviews, Surveys, Observation) +1
- 67%
Psychometric Properties of the User Experience Questionnaire (UEQ)
CHI '22· User Research Methods (Interviews, Surveys, Observation) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)