Robust Methods for Developer Screening in Rapidly Evolving AI Contexts
Authors
Paper Title
Robust Methods for Developer Screening in Rapidly Evolving AI Contexts
Publication Info
- Topic area: Screening methods for identifying programmers in empirical studies resistant to AI-assisted cheating.
- Keywords: AI resistance, ChatGPT, screening questions, programmer identification, audio-based tasks, empirical research, usable security, software engineering, LLMs, time constraints.
Background and Problem
- Problem / challenge: Existing screening methods for identifying programmers in empirical studies are vulnerable to AI-assisted cheating, particularly from tools like ChatGPT. Static visual formats and brief time limits can be circumvented using advanced AI capabilities.
- Significance: Ensuring the internal validity of studies that rely on programmer expertise is critical, as the inclusion of non-programmers can distort findings and undermine generalizability.
- Motivation and related work: Prior work proposed static visual and code comprehension tasks as ChatGPT-resistant screeners but found them increasingly ineffective against AI advancements. This paper builds on these efforts by exploring dynamic and sequential formats that exploit AI limitations in real-time processing.
Solution
- Proposed approach: Sequential audio-based screening questions designed to resist AI-assisted cheating by presenting programming-related terms with strict time constraints.
- Novelty:
- Development of six audio-based screening questions that reliably differentiate programmers from non-programmers, even with AI assistance.
- Introduction of sequential formats that reveal information step by step, minimizing opportunities for AI tools to process and respond within time limits.
- Empirical evaluation demonstrating robustness against ChatGPT and providing configuration guidelines for effective screening setups.
- Procedure and key techniques:
- Design of 28 screening questions (12 video-based, 16 audio-based) with sequential formats.
- Implementation of strict time limits and gradual information revelation to hinder AI-assisted responses.
- Recruitment and testing with 74 participants divided into three groups: programmers without ChatGPT, non-programmers without ChatGPT, and non-programmers using ChatGPT.
- Statistical analysis to identify effective screening questions based on predefined correctness thresholds.
Results
- Concrete findings:
- Six audio-based screening questions met the criteria for effectiveness, with programmers achieving ≥92% correctness and non-programmers ≤33.34%.
- Recommended configurations allow for 95.87–99.69% programmer inclusion while excluding 99.69% of non-programmers.
- Non-programmers using ChatGPT struggled to answer correctly due to time constraints and cognitive load.
- Advantage over baselines:
- Audio-based tasks outperform static visual screeners by leveraging sequential formats and strict timing to exploit AI limitations in real-time processing.
- Programmers demonstrated high accuracy with minimal perceived time pressure, while non-programmers experienced significant difficulty.
- Experiments / evaluation:
- Study conducted online with 74 participants recruited via Upwork.
- Evaluation included correctness rates, perceived time pressure, and strategies for using ChatGPT.
- Statistical testing confirmed the robustness of recommended audio questions against AI-assisted cheating.
- Limitations and future work:
- Accessibility concerns for participants with hearing impairments or language-processing difficulties.
- Focus on ChatGPT may not generalize to other AI tools.
- Potential improvements in AI capabilities could necessitate further refinement of screening methods.
Summary
This paper presents a novel approach to screening programmers in empirical studies by introducing audio-based tasks resistant to AI-assisted cheating. By leveraging sequential formats and strict time limits, the proposed method effectively distinguishes programmers from non-programmers, even when AI tools like ChatGPT are used. The study identifies six robust screening questions and provides configuration guidelines for efficient and reliable implementation. While accessibility and future AI advancements remain challenges, this work offers a practical solution for maintaining internal validity in programmer studies and highlights broader implications for online research in the age of LLMs.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Questioning the AI: Informing Design Practices for Explainable AI User Experiences
CHI '20· Explainable AI (XAI) +1
- 83%
How can Explainability Methods be Used to Support Bug Identification in Computer Vision Models?
CHI '22· Explainable AI (XAI) +1
- 83%
Zeno: An Interactive Framework for Behavioral Evaluation of Machine Learning
CHI '23· Explainable AI (XAI) +1
- 71%
Adapting User Interfaces with Model-based Reinforcement Learning
CHI '21· Human-LLM Collaboration +2
- 71%
RiskRAG: A Data-Driven Solution for Improved AI Model Risk Reporting
CHI '25· Explainable AI (XAI) +2
- 71%
A decision-theoretic representation of assistive interfaces
CHI '26· AI-Assisted Decision-Making & Automation +2
- 71%
The AI Memory Gap: Users Misremember What They Created With AI or Without
CHI '26· Human-LLM Collaboration +2
- 71%
Uncovering Relationships Between Android Developers, User Privacy, and Developer Willingness to Reduce Fingerprinting Risks
CHI '26· Privacy by Design & User Control +2
- 71%
Keeping Designers in the Loop: Communicating Inherent Algorithmic Trade-offs Across Multiple Objectives
DIS '20· Explainable AI (XAI) +2
- 67%
UMLAUT: Debugging Deep Learning Programs using Program Structure and Model Behavior
CHI '21· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)