Small, Medium, Large? A Meta-Study of Effect Sizes at CHI to Aid Interpretation of Effect Sizes and Power Calculation
Authors
Research Background and Issues
- Issues and Challenges: This study identifies a low frequency of reporting and discussion of statistical results, particularly effect sizes, in quantitative research within the CHI community. These issues significantly impact the evaluation of practical significance, prospective power analysis, and the integration of research findings within the field. Furthermore, most interpretations are limited to irrelevant external structures, with a lack of specific guidelines for effect size judgment.
- Significance: Effect sizes facilitate the assessment of practical impact, provide scientifically valid guidance for sample size calculation, and reduce the risks of research redundancy, experimental design inefficiencies, and resource waste due to the absence of field-specific effect size standards.
- Research Motivation and Related Work: Previous studies in fields such as psychology and medicine have emphasized the importance of extracting meaningful information from statistical reports. However, within HCI, particularly the CHI community, discussions of effect sizes are notably lacking. This study aims to address this gap by developing tools and constructing effect size distributions tailored to the CHI domain, providing guidance for the community.
Solution
- Proposed Solution:
- Developed a statistical information extraction tool based on large language models (GPT-4) to extract and classify statistical data from CHI papers.
- Conducted a comprehensive quantitative analysis of CHI papers from 2019 to 2023, summarizing effect size distributions across 12 research domains.
- Proposed an analytical framework for effect sizes and identified five primary methods for interpreting effect sizes in CHI research.
- Innovations:
- Developed a novel extraction tool for automated data analysis, incorporating consistency checks to reduce model hallucinations.
- Constructed domain-specific effect size thresholds within the CHI community based on extracted statistical data, addressing the lack of field-specific guidance.
- Conducted qualitative analysis of effect size interpretation in CHI papers and proposed a structured framework for discussing effect sizes.
- Implementation Steps and Techniques:
- Retrieved and filtered quantitative CHI papers.
- Used GPT-4 to extract basic statistical test results, including p-values, effect sizes, confidence intervals, and sample sizes.
- Unified various types of effect sizes into Pearson's r and divided effect size thresholds using critical points.
- Conducted qualitative analysis to summarize methods and practices for interpreting reported effect sizes.
Research Outcomes
- Specific Outcomes:
- Extracted over 6,000 effect sizes from 1,692 quantitative CHI papers and established effect size threshold guidelines across 12 research domains.
- Proposed five classification methods for interpreting effect sizes: comparing test values (e.g., describing direction), assigning size labels, referencing statistical frameworks, comparing different observations, and "big picture" explanations.
- Advantages:
- Provided domain-specific effect size references for 12 CHI fields, aiding researchers in better understanding results and conducting prospective power analyses.
- Combined quantitative and qualitative analyses to optimize effect size discussions, encouraging transparent reporting and practical interpretations.
- Experimental or Evaluation Results:
- Testing the tool's extraction performance revealed that all effect sizes were correctly extracted, though some effect sizes were missed (approximately 32%).
- The study demonstrated significant differences in overall effect size distributions across CHI, with clear variations in thresholds related to specific research domains.
- Analysis of 67 CHI'23 quantitative papers indicated that the use of effect size reporting still requires improvement, particularly in domain-specific interpretations.
- Limitations and Future Directions:
- The extraction tool has limitations in false negatives, leading to data omissions.
- Errors may occur when introducing correction factors during effect size conversion, especially in comparing effects between different design modes (between-group vs. within-group).
- Future work should focus on further automating report generation, embedding statistical test information into supplementary materials, and expanding guidelines to cover additional effect metrics.
Conclusion
This study proposes an effect size reporting and interpretation framework tailored to the CHI community through empirical analysis and tool development, providing comprehensive guidance for researchers in effect size evaluation and power analysis. Future efforts could expand automated tools for effect size standardization and develop practical methods for broader domains. This work offers significant contributions to improving the quality and transparency of quantitative research in CHI.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Why is effect size reporting and discussion infrequent in CHI quantitative research, and how does this affect synthesis and practical evaluation in the field?Category: Research Synthesis, Domain Reflection, and Methodological PerspectivesSimilar questionsarrow_forward
- How can tools and methods be developed to define and promote interpretation and standards for specific effect sizes in CHI?Category: Research Synthesis, Domain Reflection, and Methodological PerspectivesSimilar questionsarrow_forward
- How should effect size threshold ranges and classification methods be designed in CHI research to aid interpretation and power analysis?Category: Research Synthesis, Domain Reflection, and Methodological PerspectivesSimilar questionsarrow_forward
Practical Problems
1- CHI researchers struggle to assess practical significance of statistical results, leading to wasted resources and low design efficiency.Category: Research Synthesis, Domain Reflection, and Methodological PerspectivesSimilar questionsarrow_forward
- 83%
RDoFlow: Automatically assessing under-specified statistical analyses in HCI
IUI '26· User Research Methods (Interviews, Surveys, Observation) +2
- 80%
A Qualitative Study on How Usable Security and HCI Researchers Judge the Size and Importance of Odds Ratio and Cohen's d Effect Sizes
CHI '25· User Research Methods (Interviews, Surveys, Observation) +1
- 80%
Better Assumptions, Stronger Conclusions: The Case for Ordinal Regression in HCI
CHI '26· User Research Methods (Interviews, Surveys, Observation) +2
- 80%
On the Computational Reproducibility of Human-Computer Interaction
CHI '26· Research Ethics & Open Science +2
- 80%
Opportunities and Limits of Online Video Studies
DIS '25· User Research Methods (Interviews, Surveys, Observation) +2
- 80%
An Aligned Rank Transform Procedure for Multifactor Contrast Tests
UIST '21· User Research Methods (Interviews, Surveys, Observation) +1
- 67%
Special Interest Group on Transparent Statistics Guidelines
CHI '18· Computational Methods in HCI +1
- 67%
Designing for Reproducibility: A Qualitative Study of Challenges and Opportunities in High Energy Physics
CHI '19· User Research Methods (Interviews, Surveys, Observation) +2
- 60%
HARK No More: On the Preregistration of CHI Experiments
CHI '18· Computational Methods in HCI +1
- 60%
Standardizing Participant Compensation Reporting in HCI: A Meta-Review and Recommendations for the Field
CHI '21· User Research Methods (Interviews, Surveys, Observation) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)