Small, Medium, Large? A Meta-Study of Effect Sizes at CHI to Aid Interpretation of Effect Sizes and Power Calculation

User Research Methods (Interviews, Surveys, Observation)Computational Methods in HCIResearch Ethics & Open ScienceHCI ResearchersStatisticians & Data Scientists

Research Background and Issues

  • Issues and Challenges: This study identifies a low frequency of reporting and discussion of statistical results, particularly effect sizes, in quantitative research within the CHI community. These issues significantly impact the evaluation of practical significance, prospective power analysis, and the integration of research findings within the field. Furthermore, most interpretations are limited to irrelevant external structures, with a lack of specific guidelines for effect size judgment.
  • Significance: Effect sizes facilitate the assessment of practical impact, provide scientifically valid guidance for sample size calculation, and reduce the risks of research redundancy, experimental design inefficiencies, and resource waste due to the absence of field-specific effect size standards.
  • Research Motivation and Related Work: Previous studies in fields such as psychology and medicine have emphasized the importance of extracting meaningful information from statistical reports. However, within HCI, particularly the CHI community, discussions of effect sizes are notably lacking. This study aims to address this gap by developing tools and constructing effect size distributions tailored to the CHI domain, providing guidance for the community.

Solution

  • Proposed Solution:
    • Developed a statistical information extraction tool based on large language models (GPT-4) to extract and classify statistical data from CHI papers.
    • Conducted a comprehensive quantitative analysis of CHI papers from 2019 to 2023, summarizing effect size distributions across 12 research domains.
    • Proposed an analytical framework for effect sizes and identified five primary methods for interpreting effect sizes in CHI research.
  • Innovations:
    • Developed a novel extraction tool for automated data analysis, incorporating consistency checks to reduce model hallucinations.
    • Constructed domain-specific effect size thresholds within the CHI community based on extracted statistical data, addressing the lack of field-specific guidance.
    • Conducted qualitative analysis of effect size interpretation in CHI papers and proposed a structured framework for discussing effect sizes.
  • Implementation Steps and Techniques:
    • Retrieved and filtered quantitative CHI papers.
    • Used GPT-4 to extract basic statistical test results, including p-values, effect sizes, confidence intervals, and sample sizes.
    • Unified various types of effect sizes into Pearson's r and divided effect size thresholds using critical points.
    • Conducted qualitative analysis to summarize methods and practices for interpreting reported effect sizes.

Research Outcomes

  • Specific Outcomes:
    • Extracted over 6,000 effect sizes from 1,692 quantitative CHI papers and established effect size threshold guidelines across 12 research domains.
    • Proposed five classification methods for interpreting effect sizes: comparing test values (e.g., describing direction), assigning size labels, referencing statistical frameworks, comparing different observations, and "big picture" explanations.
  • Advantages:
    • Provided domain-specific effect size references for 12 CHI fields, aiding researchers in better understanding results and conducting prospective power analyses.
    • Combined quantitative and qualitative analyses to optimize effect size discussions, encouraging transparent reporting and practical interpretations.
  • Experimental or Evaluation Results:
    • Testing the tool's extraction performance revealed that all effect sizes were correctly extracted, though some effect sizes were missed (approximately 32%).
    • The study demonstrated significant differences in overall effect size distributions across CHI, with clear variations in thresholds related to specific research domains.
    • Analysis of 67 CHI'23 quantitative papers indicated that the use of effect size reporting still requires improvement, particularly in domain-specific interpretations.
  • Limitations and Future Directions:
    • The extraction tool has limitations in false negatives, leading to data omissions.
    • Errors may occur when introducing correction factors during effect size conversion, especially in comparing effects between different design modes (between-group vs. within-group).
    • Future work should focus on further automating report generation, embedding statistical test information into supplementary materials, and expanding guidelines to cover additional effect metrics.

Conclusion

This study proposes an effect size reporting and interpretation framework tailored to the CHI community through empirical analysis and tool development, providing comprehensive guidance for researchers in effect size evaluation and power analysis. Future efforts could expand automated tools for effect size standardization and develop practical methods for broader domains. This work offers significant contributions to improving the quality and transparency of quantitative research in CHI.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188667/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713671
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
User Research Methods (Interviews, Surveys, Observation), Computational Methods in HCI, Research Ethics & Open Science
work
Professions
HCI Researchers, Statisticians & Data Scientists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers