What Happens When Reviewers Receive AI Feedback in Their Reviews?
Authors
Paper Title
What Happens When Reviewers Receive AI Feedback in Their Reviews?
Publication Info
- Topic area: The role and impact of AI feedback tools in academic peer review processes.
- Keywords: AI feedback, peer review, ICLR 2025, reviewer perceptions, academic publishing, human-AI collaboration, review quality, governance, mixed-methods study, academic sustainability.
Background and Problem
- Problem / challenge: Rising submission volumes in academic conferences are straining the peer review system, leading to inconsistent review quality and reviewer fatigue. Existing AI tools for peer review are underexplored, particularly in providing post-review feedback to reviewers.
- Significance: Understanding how reviewers perceive and respond to AI feedback is critical for designing tools that enhance review quality while safeguarding human expertise and accountability.
- Motivation and related work: Prior research has focused on hypothetical AI scenarios or automation in reviewing. Little is known about how reviewers engage with AI feedback in live, high-stakes settings. This study builds on prior work by examining the real-world deployment of an AI feedback tool at ICLR 2025.
Solution
- Proposed approach: Deployment and study of an AI feedback tool at ICLR 2025, which provided post-review suggestions to reviewers to improve clarity, tone, and specificity.
- Novelty:
- First empirical study of AI feedback in a live, high-stakes peer review process.
- Insights into human–AI negotiation in peer review, highlighting tensions between support and evaluative authority.
- Design implications for AI tools that enhance review quality while preserving human agency and responsibility.
- Procedure and key techniques:
- Mixed-methods study combining surveys (N=51) and interviews (N=9) with ICLR 2025 reviewers.
- Analysis of reviewers’ perceptions, actions, and envisioned roles for AI in peer review.
- Thematic analysis of qualitative data and statistical analysis of survey responses.
Results
- Concrete findings:
- 78.4% of survey participants considered revising their reviews after receiving AI feedback, but only 56.9% followed through.
- AI feedback was perceived as relevant (M = 4.51, p = 0.032) and actionable (M = 4.22), but less constructive (M = 3.88) and occasionally inappropriate (M = 4.39).
- Reviewers did not feel that AI reduced their sense of ownership (M = 3.94) or accountability (M = 3.63).
- Advantage over baselines: The study moves beyond speculative or simulated settings by analyzing real reviewer interactions with an AI tool, providing actionable insights into its practical impact.
- Experiments / evaluation:
- Survey: Likert-scale and open-ended questions to assess perceptions of usefulness, feedback quality, and behavioral intentions.
- Interviews: Semi-structured discussions to explore reviewers’ experiences, concerns, and visions for future AI tools.
- Participants: 51 survey respondents and 9 interviewees, primarily early-career researchers with diverse reviewing experience.
- Limitations and future work:
- Sample skewed toward male and early-career researchers, limiting diversity.
- Reliance on self-reported data rather than objective measures of review quality.
- Findings specific to ICLR 2025; further research needed across different venues and review structures.
Summary
This study investigates the impact of AI feedback on reviewers during the ICLR 2025 review process. While reviewers acknowledged that the AI tool improved clarity and specificity, they often perceived it as less useful because it focused on surface-level expression rather than intellectual evaluation. The study highlights a paradox: reviewers adopted AI suggestions but resisted its authority, emphasizing the need for tools that align with their intellectual challenges. Participants envisioned AI as a collaborator in reasoning, a labor-relief tool, and a quality safeguard, while stressing the importance of human evaluative judgment. These findings provide actionable insights for designing AI tools that enhance review quality without undermining reviewer autonomy.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and Considerations
CHI '26· Human-LLM Collaboration +2
- 80%
Transparency of CHI Research Artifacts: Results of a Self-Reported Survey
CHI '20· Explainable AI (XAI) +1
- 71%
All Accept, No Reject: Evaluating LLMs as “Peer” Reviewers
CHI '26· Human-LLM Collaboration +3
- 67%
Integrating measures of replicability into scholarly search: Challenges and opportunities
CHI '24· Explainable AI (XAI) +2
- 67%
Large Language Models in Qualitative Research: Uses, Tensions, and Intentions
CHI '25· Human-LLM Collaboration +2
- 60%
Changes in Research Ethics, Openness, and Transparency in Empirical Studies between CHI 2017 and CHI 2022
CHI '23· Research Ethics & Open Science
Based on Jaccard similarity of research subtopics & professions (≥60%)