Quantifying the Novelty Bias when Evaluating Interactive Prototypes

Honorable Mention
User Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingResearch Ethics & Open ScienceHCI ResearchersCognitive ScientistsUI/UX Designers

Paper Title

Quantifying the Novelty Bias when Evaluating Interactive Prototypes

Publication Info

  • Topic area: Human-computer interaction (HCI) evaluation and novelty bias.
  • Keywords: Novelty effect, human-computer interaction, subjective ratings, performance metrics, technology readiness, prototype evaluation, placebo effect, user studies, bias quantification, interaction design.

Background and Problem

  • Problem / challenge: HCI evaluations often rely on user judgments to assess new technologies, but novelty bias—favoring what is labeled as “new”—can distort these judgments, leading to inflated preferences and ratings that may not align with actual performance.
  • Significance: Understanding and quantifying novelty bias is crucial for ensuring the validity of HCI evaluations, particularly as they inform design decisions and claims of improvement.
  • Motivation and related work: Prior research has acknowledged novelty effects in HCI but has not quantified their impact on subjective ratings and objective performance. Related studies on placebo effects and evaluation validity highlight the need to isolate novelty framing as a distinct biasing force.

Solution

  • Proposed approach: A controlled experimental study to quantify the impact of novelty labeling on subjective preference and objective performance across four technology types: mice, keyboards, search engines, and AI chatbots.
  • Novelty:
    1. Demonstrates causally that labeling a functionally identical system as “new” shifts preferences and subjective ratings.
    2. Shows that these perceptual shifts can outweigh or misalign with objective performance metrics.
    3. Clarifies the role of technology readiness (TR) in moderating baseline performance but not shielding judgments from novelty bias.
  • Procedure and key techniques:
    • Participants interacted with functionally identical technology pairs differing only in cosmetic features and “old”/“new” labels.
    • Tasks included pointing, typing, searching, and chatting, with measures of performance (e.g., throughput, error rates) and subjective ratings (e.g., Likert scales).
    • Extensive counterbalancing ensured no systematic bias from order effects or cosmetic preferences.
    • Statistical models analyzed trial-level and participant-level outcomes, including interactions with TR.

Results

  • Concrete findings:
    • Up to 77% of participants preferred the “new” versions, with subjective ratings inflated by up to 11.6% for preferred AI chatbots and 7.1% for the “new” search engine.
    • Performance differences were modest: 9.7% fewer misses with the “new” mouse and up to 7.2% lower error rates with the “new” keyboard.
    • Throughput, movement time, browsing time, and number of messages exchanged showed little change.
  • Advantage over baselines: Novelty labeling consistently shifted preferences and ratings, but objective performance metrics showed limited improvement, highlighting a misalignment between perceived and actual benefits.
  • Experiments / evaluation:
    • Sample size: 48 participants across diverse backgrounds.
    • Technologies tested: mice, keyboards, search engines, AI chatbots.
    • Metrics: speed, accuracy, subjective ratings, and TR scores.
    • Statistical methods: mixed-effects models, nonparametric tests, and binomial tests for preference.
  • Limitations and future work:
    • Results focus on first-use, structured lab settings; longer-term and field studies are needed to assess novelty effects over time.
    • Manipulation relied on explicit labels; future work could explore subtler novelty cues.
    • Broader technology classes and higher-level tasks (e.g., collaboration, creative work) may reveal different sensitivities to novelty.

Summary

This study quantifies the impact of novelty labeling on subjective preference and objective performance across four technology types. Results show that novelty framing strongly influences preference and subjective ratings, with limited effects on performance metrics such as speed and accuracy. Technology readiness predicts baseline skill but does not mitigate novelty bias. These findings highlight the methodological challenge of separating genuine improvements from novelty-driven perceptions in HCI evaluations. Future research should explore longitudinal effects, subtler novelty cues, and broader technology domains to refine evaluation practices and ensure reliable conclusions about system quality.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222350/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791685
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
3 authors
sell
Subtopics
User Research Methods (Interviews, Surveys, Observation), Prototyping & User Testing, Research Ethics & Open Science
work
Professions
HCI Researchers, Cognitive Scientists, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers