Wikipedia ORES Explorer: Visualizing Trade-offs For Designing Applications With Machine Learning API
Authors
With the growing industry applications of Artificial Intelligence (AI) systems, pre-trained models and APIs have emerged and greatly lowered the barrier of building AI-powered products. However, novice AI application designers often struggle to recognize the inherent algorithmic trade-offs and evaluate model fairness before making informed design decisions. In this study, we examined the Objective Revision Evaluation System (ORES), a machine learning (ML) API in Wikipedia used by the community to build anti-vandalism tools. We designed an interactive visualization system to communicate model threshold trade-offs and fairness in ORES. We evaluated our system by conducting 10 in-depth interviews with potential ORES application designers. We found that our system helped application designers who have limited ML backgrounds learn about in-context ML knowledge, recognize inherent value trade-offs, and make design decisions that aligned with their goals. By demonstrating our system in a real-world domain, this paper presents a novel visualization approach to facilitate greater accessibility and human agency in AI application design.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Lexara: A User-Centered Toolkit for Evaluating Large Language Models for Conversational Visual Analytics
CHI '26· Human-LLM Collaboration +2
- 67%
The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With Reality
CHI '21· Explainable AI (XAI) +1
- 67%
Angler: Helping Machine Translation Practitioners Prioritize Model Improvements
CHI '23· Explainable AI (XAI) +2
- 67%
"Are You Really Sure?'' Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making
CHI '24· Explainable AI (XAI) +1
- 67%
Quantitative Evaluation of Machine Learning Explanations: A Human-Grounded Benchmark
IUI '21· Explainable AI (XAI) +1
- 67%
iSEA : An Interactive Pipeline for Semantic Error Analysis of NLP Models
IUI '22· Explainable AI (XAI) +2
- 67%
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
IUI '25· Explainable AI (XAI) +1
- 67%
ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts
UIST '25· Human-LLM Collaboration +2
- 60%
Conversational Explanations: Discussing Explainable AI with Non-AI Experts
IUI '25· Explainable AI (XAI)
Based on Jaccard similarity of research subtopics & professions (≥60%)