Evaluation-First Design for Data Visualization Interfaces
Authors
Paper Title
Evaluation-First Design for Data Visualization Interfaces
Publication Info
- Topic area: Data visualization and human-computer interaction (HCI) design methodologies.
- Keywords: Evaluation-first design, data visualization, design study methodology, stakeholder engagement, iterative evaluation, metrics, feedback loops, co-evaluation, design frameworks, visualization tools.
Background and Problem
- Problem / challenge: Existing visualization design frameworks, such as DSM and Data-First, do not fully integrate evaluation as a continuous, explicit component across all design phases. They lack guidance on evaluation timing, role allocation, and how evaluation evidence evolves and persists.
- Significance: Addressing these gaps can improve alignment, trust, and decision-making in visualization projects, especially under time constraints or with evolving requirements.
- Motivation and related work: Prior frameworks like DSM and Data-First acknowledge evaluation but treat it as implicit and cross-cutting rather than central. HCI research has explored evaluation types but not its role in design processes. This paper builds on these foundations to propose a more explicit and structured approach to evaluation.
Solution
- Proposed approach: Evaluation-first design (EvalOps), an overlay to existing frameworks that integrates evaluation as a continuous and explicit component across all design phases.
- Novelty:
- Introduces tighter feedback loops (small and large) to ensure evaluation is actionable and traceable.
- Establishes co-evaluation with stakeholders, making them active participants in shaping and interpreting evaluation.
- Implements goals-to-metrics grounding, linking stakeholder goals to evolving evaluation metrics.
- Procedure and key techniques:
- Cadence: Nested feedback loops (small, informal checks and larger, structured reviews) ensure continuous evaluation.
- Roles: Structured participation, where stakeholders act as co-evaluators alongside designers and developers.
- Metrics: Goals-to-metrics grounding aligns evaluation activities with stakeholder-defined success measures, evolving as the project progresses.
Results
- Concrete findings:
- EvalOps enabled timely pivots and evidence-backed decisions in two case studies, improving tool usability and alignment with stakeholder needs.
- Frequent loop closures and structured participation ensured evaluation evidence was actionable and traceable.
- Advantage over baselines:
- Compared to DSM and Data-First, EvalOps explicitly integrates evaluation across all phases, making it continuous, role-structured, and metric-grounded.
- Enabled faster course corrections and better alignment with stakeholder priorities.
- Experiments / evaluation:
- Two case studies: (1) AI-enabled summary-shaping tool and (2) AI-enabled speech-to-text (STT) triage tool.
- Evaluation included small loops (e.g., weekly prototype reviews) and large loops (e.g., quarterly deployments or monthly reviews).
- Metrics evolved from baseline measures (e.g., perceived summary quality) to task-specific success signals (e.g., timeline reconstruction accuracy).
- Limitations and future work:
- Limited generalization due to only two case studies; further validation across domains is needed.
- Stakeholder participation may not always be feasible, requiring adaptations for limited access.
- Metrics must remain flexible, avoiding early fixation. Future work includes refining evaluation-first design sheets and validating the approach in broader contexts.
Summary
The paper introduces EvalOps, an evaluation-first overlay for visualization design frameworks like DSM and Data-First. By emphasizing cadence, roles, and metrics, EvalOps makes evaluation explicit, continuous, and actionable across all design phases. Case studies of AI-enabled tools demonstrate how EvalOps supports timely pivots, stakeholder co-evaluation, and evolving metrics, resulting in better-aligned and more usable tools. While promising, further validation across diverse domains and teams is needed to generalize the approach and refine its practices.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
Perceptual Pat: A Virtual Human Visual System for Iterative Visualization Design
CHI '23· Interactive Data Visualization +2
- 75%
SemTabla: A Human-in-the-Loop Framework for Semantic Enrichment and Validation of Data Tables
CHI '26· Explainable AI (XAI) +3
- 71%
Does a Picture Paint a Thousand Words? Using Visual and Textual Channels to Understand Attitudes and Beliefs
CHI '26· User Research Methods (Interviews, Surveys, Observation) +2
- 63%
"I Need to Find That One Chart": How Data Workers Navigate, Summarize and Communicate Analytical Conversations
CHI '26· User Research Methods (Interviews, Surveys, Observation) +2
- 63%
Interface Dis/Similarities: Investigating Characteristics Influencing Perceived Differences Between GUIs
CHI '26· User Research Methods (Interviews, Surveys, Observation) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)