Chasing Meaning and/or Insight? A Survey on Evaluation Practices at the Intersection of Visualization and the Humanities
Best PaperAuthors
National Distance Education University
University for Continuing Education Krems
University of Luxembourg
Chemnitz University of Technology
King's College London
University for Continuing Education Krems
Paper Title
Chasing Meaning and/or Insight? A Survey on Evaluation Practices at the Intersection of Visualization and the Humanities
Publication Info
- Topic area: Evaluation practices in visualization for humanities research.
- Keywords: Visualization, humanities, evaluation, interpretive methods, digital humanities, user-centered design, mixed methods, provenance, uncertainty, epistemology.
Background and Problem
- Problem / challenge: Current evaluation practices in visualization for humanities research (VIS*H) often rely on monomethod approaches that fail to rigorously validate interpretive depth and scholarly utility. There is a mismatch between the positivist evaluation frameworks of visualization and the interpretive aims of humanities scholarship.
- Significance: Addressing this gap is critical for advancing visualization tools that genuinely support humanities inquiry, bridging the divide between functional performance and interpretive power.
- Motivation and related work: Previous studies have documented visualization methods and challenges in digital humanities, but comprehensive surveys of evaluation practices specific to VIS*H are lacking. This paper builds on foundational work in visualization evaluation and humanities-centered design to analyze methodological gaps and propose improvements.
Solution
- Proposed approach: A systematic survey of 171 VIS*H design studies to analyze evaluation workflows, identify methodological gaps, and derive recommendations for improving evaluation rigor and alignment with humanities practices.
- Novelty:
- Comprehensive coding framework capturing evaluation methods, study designs, participant profiles, and visualization features.
- Statistical analysis of evaluation quality, identifying methods and workflows that predict higher rigor.
- Identification of eight archetypical evaluation workflows through cluster analysis.
- Evidence-based recommendations for improving VIS*H evaluation practices.
- Procedure and key techniques:
- Development of a seed corpus of seminal VIS*H publications.
- Snowball sampling to expand the dataset to 1,317 papers, filtered to 171 design studies meeting inclusion criteria.
- Coding of papers across multiple dimensions, including evaluation methods, participant demographics, and study designs.
- Statistical modeling (ordinal logistic regression) to assess predictors of evaluation quality.
- Hierarchical cluster analysis to identify recurrent methodological patterns.
Results
- Concrete findings:
- Mean evaluation quality score across studies was 1.98 on a 0–5 scale, with 60% of papers scoring 2 or below.
- Log analysis (OR ≈ 4.86), questionnaires/surveys (OR ≈ 4.03), and interviews (OR ≈ 3.21) were the strongest predictors of higher evaluation quality.
- Mixed-method designs combining qualitative and quantitative approaches achieved the highest rigor.
- Only 20% of papers explicitly discussed evaluation challenges, but these had significantly higher quality scores.
- Advantage over baselines:
- Papers employing diverse evaluation methods, particularly evidence-rich techniques like log analysis and structured surveys, demonstrated significantly higher rigor compared to monomethod approaches.
- Experiments / evaluation:
- Analysis of 171 papers revealed eight archetypical evaluation workflows, ranging from low-rigor monomethod designs (e.g., case studies alone) to high-rigor mixed-method approaches (e.g., log analysis paired with interviews and testing).
- Participant gaps were identified: while 87% of tools targeted experts, only 31% were evaluated with them.
- Limitations and future work:
- The survey did not capture non-peer-reviewed formats like books or workshop papers, potentially missing innovative practices.
- Evaluation quality scoring was subjective and shaped by the authors' interdisciplinary perspectives.
- Future work should focus on developing comprehensive frameworks for validating interpretive and epistemological aspects of VIS*H tools.
Summary
This paper systematically surveys evaluation practices in visualization for humanities research, identifying methodological gaps and proposing evidence-based improvements. Key findings include the predominance of low-rigor monomethod evaluations and the importance of mixed-method designs for achieving higher quality. Statistical analysis highlights log analysis, structured surveys, and interviews as critical components of rigorous evaluations. Recommendations emphasize triangulating methods, recruiting valid participants, and integrating formative evaluation throughout the design process. Future challenges include grounding visualizations in provenance, uncertainty, and humanities theories to better align with interpretive and critical traditions. The study provides actionable insights for advancing evaluation practices in the VIS*H field.