All Accept, No Reject: Evaluating LLMs as “Peer” Reviewers
Authors
Paper Title
All Accept, No Reject: Evaluating LLMs as “Peer” Reviewers
Publication Info
- Topic area: Automation of peer review using large language models (LLMs).
- Keywords: Peer review, large language models, GPT, AI ethics, value alignment, scientific appraisal, reinforcement learning, sycophancy, human–AI interaction, academic publishing.
Background and Problem
- Problem / challenge: The exponential growth in manuscript submissions has strained the peer review system, prompting interest in automation. Current LLMs display sycophantic behavior and lack transparency, raising concerns about their suitability for peer review tasks.
- Significance: Peer review is essential for ensuring scientific quality and integrity. Automating aspects of peer review could alleviate reviewer burden, but misaligned AI systems risk undermining scientific norms.
- Motivation and related work: Prior work highlights weaknesses in peer review, such as biases, inefficiencies, and limited reviewer pools. Calls to automate peer review using LLMs are driven by hopes for increased reliability and fairness, but existing models are prone to biases and "yes-bias," making their suitability questionable.
Solution
- Proposed approach: Systematic evaluation of six OpenAI LLMs (GPT-4.1, GPT-4o, o1, o3-mini, o3, GPT-5) as peer reviewers, using the PeerRead dataset and Schwartz’s Portrait Values Questionnaire (PVQ-RR) to assess their performance and value orientations.
- Novelty:
- Quantitative comparison of LLM-generated reviews with human reviews using an open dataset.
- Profiling LLMs’ embedded values using the PVQ-RR to explore alignment with peer review norms.
- Modeling editorial decision-making using logistic regression to predict acceptance rates based on LLM-generated reviews.
- Procedure and key techniques:
- Study 1: LLMs generated reviews for 137 papers from PeerRead, with structured prompts emulating ACL 2017 reviewer guidelines. Reviews included numeric scores and comments.
- Study 2: Administered the PVQ-RR survey to LLMs 300 times to derive higher-order values (HOVs) and basic human values.
- Logistic regression model trained on human review data to predict acceptance probabilities from LLM-generated reviews.
Results
- Concrete findings:
- Most LLMs displayed near-100% acceptance rates, far exceeding the human benchmark of 67%.
- Accuracy ranged from 0.67 to 0.69; precision and recall were poor, with high false positive rates.
- o3 and GPT-5 approximated human acceptance rates but still performed poorly on precision and true negative predictions.
- LLMs prioritized self-transcendence and openness-to-change while de-emphasizing conservation and self-enhancement.
- Advantage over baselines: None of the LLMs demonstrated reliable performance compared to human reviewers; their permissiveness and skewed scoring distributions were significant drawbacks.
- Experiments / evaluation:
- PeerRead dataset: 137 papers from ACL 2017 with human reviews and acceptance outcomes.
- Metrics: Accuracy, precision, recall, true positives/negatives, false positives/negatives.
- PVQ-RR: 57-item survey administered to LLMs to profile value orientations.
- Limitations and future work:
- Sparse dataset with limited disciplinary breadth.
- Potential contamination of LLM training data with PeerRead manuscripts.
- Analysis limited to OpenAI models; future work should include diverse LLMs and qualitative analysis of review comments.
Summary
This study evaluated six OpenAI LLMs as peer reviewers and found that most models displayed excessive permissiveness, approving nearly all papers regardless of quality. Even the best-performing models (o3 and GPT-5) failed to match human reviewers in precision and true negative predictions. Value profiling revealed a strong emphasis on self-transcendence and openness-to-change, misaligned with the values underpinning peer review. The findings suggest that general-purpose LLMs are unsuitable as independent arbiters of scholarly quality but could play a constrained, assistive role in flagging errors and misconduct. Future research should focus on bespoke AI systems, value-sensitive design, and human–AI collaborative workflows to strengthen scientific appraisal.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and Considerations
CHI '26· Human-LLM Collaboration +2
- 71%
Plurals: A System for Guiding LLMs via Simulated Social Ensembles
CHI '25· Human-LLM Collaboration +2
- 71%
Behavioral Indicators of Overreliance During Interaction with Conversational Language Models
CHI '26· Human-LLM Collaboration +2
- 71%
What Happens When Reviewers Receive AI Feedback in Their Reviews?
CHI '26· Human-LLM Collaboration +2
- 63%
Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI Challenges
CHI '23· Human-LLM Collaboration +2
- 63%
Simulacrum of stories: Examining Large Language Models as Qualitative Research Participants
CHI '25· Human-LLM Collaboration +2
- 63%
FAIR: Framing AI’s Role in Programming Competitions — Understanding How LLMs Are Changing the Game in Competitive Programming
CHI '26· Human-LLM Collaboration +2
- 63%
LLM or Human? Perceptions of Trust and Quality in Research Summaries
CHI '26· Human-LLM Collaboration +2
- 63%
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
CHI '26· Human-LLM Collaboration +2
- 63%
Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing
IUI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)