Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and Considerations
Authors
Paper Title
Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and Considerations
Publication Info
- Topic area: Reporting and peer review challenges for LLM-integrated systems in HCI research.
- Keywords: HCI, LLM-integrated systems, peer review, reporting standards, AI uncertainty, technical evaluation, transparency, reproducibility, ethical concerns, interdisciplinary clashes.
Background and Problem
- Problem / challenge: The rise of LLM-integrated systems in HCI research has created uncertainty in reporting and reviewing standards, leading to inconsistent and often skeptical peer review practices.
- Significance: Addressing these challenges is crucial for ensuring fair evaluation, fostering innovation, and maintaining trust between authors and reviewers in HCI research.
- Motivation and related work: Prior work has highlighted issues such as the lack of reporting guidelines, ethical concerns with proprietary models, and tensions between HCI and ML/NLP reviewing norms. This paper builds on these insights by investigating authors’ and reviewers’ experiences and proposing considerations for improving reporting and reviewing practices.
Solution
- Proposed approach: A set of considerations for authors, reviewers, and the HCI community to navigate the challenges of reporting and reviewing LLM-integrated systems.
- Novelty:
- Identification of unique challenges posed by LLM uncertainty in HCI systems research.
- Empirical insights into authors’ strategies for reporting and reviewers’ expectations.
- Recommendations for balancing transparency, reproducibility, and practical constraints.
- Suggestions for fostering interdisciplinary understanding and establishing venue-specific guidelines.
- Procedure and key techniques:
- Conducted interviews with 18 authors of LLM-integrated system papers and analyzed their experiences using thematic analysis.
- Synthesized findings into actionable considerations for reporting and reviewing practices.
- Validated recommendations through feedback from six senior HCI researchers.
Results
- Concrete findings:
- Authors face heightened scrutiny and inconsistent standards in peer review, often needing to justify LLM usage and provide extensive technical evaluations.
- Reviewers demand more transparency and rigor, including reporting prompts, failure modes, and system architecture.
- Authors use strategies like de-emphasizing LLMs in framing and relying on proprietary models for practical reasons despite ethical concerns.
- Advantage over baselines:
- Provides nuanced insights into the challenges of reporting LLM-integrated systems, addressing gaps in existing guidelines.
- Proposes practical considerations tailored to HCI’s unique values and methodologies.
- Experiments / evaluation:
- Interview study with 18 participants, supplemented by reviews, rebuttals, and feedback from six senior HCI researchers.
- Analysis of reporting practices and peer review dynamics in HCI venues like CHI and UIST.
- Limitations and future work:
- Potential selection bias in participant recruitment; limited perspectives from senior scholars and industry researchers.
- Reliance on participant recall and limited primary review materials.
- Future work could explore standardized reporting frameworks and tools for LLM component engineering.
Summary
This paper investigates the challenges of reporting and reviewing LLM-integrated systems in HCI research, highlighting issues of uncertainty, transparency, and interdisciplinary clashes. Through interviews with authors and reviewers, it identifies strategies for re-establishing trust and proposes considerations for improving reporting practices and peer review standards. The findings emphasize the need for nuanced guidelines that balance transparency, reproducibility, and practical constraints, while fostering interdisciplinary understanding. These contributions aim to support the HCI community in navigating the complexities of LLM-integrated systems research.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
All Accept, No Reject: Evaluating LLMs as “Peer” Reviewers
CHI '26· Human-LLM Collaboration +3
- 83%
What Happens When Reviewers Receive AI Feedback in Their Reviews?
CHI '26· Human-LLM Collaboration +2
- 71%
LLM or Human? Perceptions of Trust and Quality in Research Summaries
CHI '26· Human-LLM Collaboration +2
- 71%
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
CHI '26· Human-LLM Collaboration +2
- 71%
Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing
IUI '26· Human-LLM Collaboration +2
- 71%
A Multimodal Investigation of Controllability and Cognitive Load in Interactive Machine Learning
IUI '26· Human-LLM Collaboration +2
- 71%
Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
IUI '26· Human-LLM Collaboration +2
- 67%
Transparency of CHI Research Artifacts: Results of a Self-Reported Survey
CHI '20· Explainable AI (XAI) +1
- 67%
User Modelling for Avoiding Overfitting in Interactive Knowledge Elicitation for Prediction
IUI '18· Human-LLM Collaboration +1
- 67%
DxHF: Providing High-Quality Human Feedback for LLM Alignment with Interactive Decomposition
UIST '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)