More than Marketing? On the Information Value of AI Benchmarks for Practitioners
Authors
Public AI benchmark results are widely broadcast by model developers as indicators of model quality within a growing and competitive market. However, these advertised scores do not necessarily reflect the traits of interest to those who will ultimately apply AI models. In this paper, we seek to understand if and how AI benchmarks are used to inform decision-making. Based on the analyses of interviews with 19 individuals who have used, or decided against using, benchmarks in their day-to-day work, we find that across these settings, participants use benchmarks as a signal of relative performance difference between models. However, whether this signal was considered a definitive sign of model superiority, sufficient for downstream decisions, varied. In academia, public benchmarks were generally viewed as suitable measures for capturing research progress. By contrast, in both product and policy, benchmarks -- even those developed internally for specific tasks -- were often found to be inadequate for informing substantive decisions. Of the benchmarks deemed unsatisfactory, respondents reported that their goals were neither well-defined nor reflective of real-world use. Based on the study results, we conclude that effective benchmarks should provide meaningful, real-world evaluations, incorporate domain expertise, and maintain transparency in scope and goals. They must capture diverse, task-relevant capabilities, be challenging enough to avoid quick saturation, and account for trade-offs in model performance rather than relying on a single score. Additionally, proprietary data collection and contamination prevention are critical for producing reliable and actionable results. By adhering to these criteria, benchmarks can move beyond mere marketing tricks into robust evaluative frameworks that accurately reflect AI progress and guide informed decision-making in both research and practical domains.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 80%
"Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI
CHI '21· Explainable AI (XAI) +1
- 80%
Fairness Evaluation in Text Classification: Machine Learning Practitioner Perspectives of Individual and Group Fairness
CHI '23· Explainable AI (XAI) +1
- 80%
Perceptions of the Fairness Impacts of Multiplicity in Machine Learning
CHI '25· Explainable AI (XAI) +1
- 67%
The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With Reality
CHI '21· Explainable AI (XAI) +1
- 67%
Beyond Expertise and Roles: A Framework to Characterize the Stakeholders of Interpretable Machine Learning and their Needs
CHI '21· Explainable AI (XAI) +2
- 67%
AI-Moderated Decision-Making: Capturing and Balancing Anchoring Bias in Sequential Decision Tasks
CHI '22· Explainable AI (XAI) +2
- 67%
Forgetting Practices in the Data Sciences
CHI '22· Explainable AI (XAI) +1
- 67%
"Are You Really Sure?'' Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making
CHI '24· Explainable AI (XAI) +1
- 67%
User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI Companions
CHI '25· Explainable AI (XAI) +2
- 67%
Wikipedia ORES Explorer: Visualizing Trade-offs For Designing Applications With Machine Learning API
DIS '21· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)