FAIR: Framing AI’s Role in Programming Competitions — Understanding How LLMs Are Changing the Game in Competitive Programming
Authors
Paper Title
FAIR: Framing AI’s Role in Programming Competitions — Understanding How LLMs Are Changing the Game in Competitive Programming
Publication Info
- Topic area: Impact of large language models (LLMs) on competitive programming workflows, fairness, and governance.
- Keywords: Competitive programming, large language models, fairness, governance, workflows, cheating, Codeforces, ICPC, AI-assisted programming.
Background and Problem
- Problem / challenge: The rise of LLMs is reshaping competitive programming, creating challenges in maintaining fairness, integrity, and credibility in contests. Prior studies have focused on LLM performance but lack insights into how human stakeholders are adapting to these shifts.
- Significance: Competitive programming serves as a key educational tool and talent pipeline for academia and industry. Ensuring fairness and credibility in contests is critical for preserving their educational and professional value.
- Motivation and related work: Previous research has explored AI use in education and coding benchmarks but has not addressed its impact on competitive programming workflows, fairness norms, or governance. This paper fills the gap by examining stakeholder responses and proposing governance measures.
Solution
- Proposed approach: A chess-inspired governance framework combining anomaly detection, expert review, community oversight, and proportional sanctions to address fairness and credibility challenges in programming contests.
- Novelty:
- Empirical insights into how LLMs are reshaping workflows across contestants, problem setters, coaches, and platform stewards.
- Analysis of fairness norms and contested boundaries between legitimate assistance and cheating.
- Proposal of governance measures inspired by mind sports like chess and Go to safeguard integrity in contests.
- Procedure and key techniques:
- Conducted 37 interviews with stakeholders (contestants, problem setters, coaches, platform stewards) and a global survey of 207 contestants.
- Analyzed platform-level data from Codeforces (336 contests, 64 million submissions) to identify behavioral proxies for AI-assisted workflows.
- Synthesized findings into actionable governance measures, including rating-linked anomaly detection, community-driven oversight, and transparent sanction ladders.
Results
- Concrete findings:
- Contestants use LLMs heavily for post-contest review and daily training but avoid them during contests due to strict rules and integrity norms.
- Problem setters integrate LLMs into low-creativity tasks like statement polishing, bilingual translation, and solver checks but exclude them from ideation.
- Coaches focus on screening for AI misuse in team selection, while platform stewards face growing burdens in monitoring and detection.
- Behavioral proxies (Python share, temporal uniformity) suggest widespread shifts toward tool-mediated workflows, even as official sanction rates remain low.
- Advantage over baselines: The proposed governance framework addresses gaps in current enforcement by combining automated detection with expert review and community oversight, ensuring fairness while minimizing false positives.
- Experiments / evaluation:
- Interviews and surveys captured stakeholder perspectives on workflows, fairness, and governance.
- Platform-level analysis provided quantitative insights into sanction rates and behavioral trends.
- Governance proposals were informed by lessons from chess and Go communities.
- Limitations and future work:
- Sample skewed toward East Asian and male participants; future studies should aim for broader demographic representation.
- Limited access to Codeforces stewards; engaging with their leadership is a priority.
- Social desirability bias may underreport cheating; future work could include social media analysis to capture emergent practices.
- Behavioral proxies are descriptive trends, not direct evidence of AI usage; longitudinal studies are needed to disentangle general ecosystem changes from AI-specific effects.
Summary
This paper investigates how LLMs are transforming competitive programming, focusing on workflows, fairness, and governance. Contestants use LLMs for training and review but avoid them during contests, while problem setters and stewards integrate them into low-creativity tasks and detection workflows. The study highlights contested fairness norms and covert AI-assisted cheating practices. Inspired by chess and Go, the authors propose a governance framework combining anomaly detection, expert review, community oversight, and proportional sanctions to safeguard integrity. Future work should expand demographic inclusivity, refine detection methods, and address the evolving role of AI in STEM competitions.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Interaction Context Often Increases Sycophancy in LLMs
CHI '26· Human-LLM Collaboration +2
- 71%
A Human-Centered Review of Algorithms for Decision-Making in Higher Education
CHI '23· AI-Assisted Decision-Making & Automation +2
- 71%
Are Two Heads Better Than One in AI-Assisted Decision Making? Comparing the Behavior and Performance of Groups and Individuals in Human-AI Collaborative Recidivism Risk Assessment
CHI '23· Human-LLM Collaboration +2
- 71%
What is Human-Centered about Human-Centered AI? A Map of the Research Landscape
CHI '23· Human-LLM Collaboration +2
- 71%
Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions
CHI '24· Human-LLM Collaboration +2
- 71%
Understanding Socio-technical Factors Configuring AI Non-Use in UX Work Practices
CHI '25· Human-LLM Collaboration +2
- 71%
When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
CHI '26· Human-LLM Collaboration +2
- 71%
Understanding Compliance and Conversion Dynamics in Multi-Agent Collectives
CHI '26· Human-LLM Collaboration +2
- 71%
Investigating the Effects of LLM Use on Critical Thinking Under Time Constraints: Access Timing and Time Availability
CHI '26· Human-LLM Collaboration +2
- 71%
Toward Scalable and Responsible Integration of Course-Specific AI Tutors: Instructor Experiences with a Campus-Wide Platform
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)