Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
Authors
Paper Title
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
Publication Info
- Topic area: Impact of AI automation on software development workflows
- Keywords: AI coding assistants, copilots, coding agents, developer productivity, user experience, software workflows, GitHub Copilot, OpenHands, human-agent interaction, LLMs
Background and Problem
- Problem / challenge: Current evaluations of coding agents rely on static benchmarks, which fail to capture real-world developer interactions. There is limited understanding of how coding agents compare to copilots in terms of productivity and user experience.
- Significance: Understanding the impact of increasingly autonomous AI tools is critical for improving developer workflows and ensuring effective adoption of these technologies.
- Motivation and related work: Prior studies have extensively analyzed copilots like GitHub Copilot, showing positive impacts on productivity. However, coding agents, which offer greater autonomy, remain underexplored in real-world settings. This paper addresses the gap by conducting a controlled study comparing copilots and agents.
Solution
- Proposed approach: A controlled user study comparing GitHub Copilot (copilot) and OpenHands (agent) to evaluate their effects on developer productivity, user experience, and interaction patterns.
- Novelty:
- First controlled study directly comparing copilots and agents in realistic coding tasks.
- Quantification of productivity improvements and user experience differences between the two tools.
- Identification of design desiderata for improving agentic workflows.
- Procedure and key techniques:
- Participants (n=20) with prior experience using GitHub Copilot but no experience with agents were recruited.
- Tasks included creating programs, adding features, and fixing bugs in Python repositories.
- Participants interacted with both tools in a within-participant setup, with randomized task order.
- Data collected included task correctness, user effort (time spent), Likert-scale survey responses, and qualitative feedback.
Results
- Concrete findings:
- Agents enabled a 35% increase in task correctness (60% vs. 25% for copilots, p=0.02).
- User effort was halved with agents (12.5 minutes vs. 25.1 minutes for copilots, p=0.01).
- Agents were particularly effective in automating environment setup and debugging tasks.
- Advantage over baselines:
- Agents outperformed copilots in enabling task completion and reducing cognitive load.
- Participants reported accomplishing new tasks with agents that were infeasible with copilots.
- Experiments / evaluation:
- Metrics: Task correctness, time spent, Likert-scale ratings, qualitative feedback.
- Tools: GitHub Copilot and OpenHands.
- Tasks: Data analysis, feature addition, and bug fixing in Python repositories.
- Limitations and future work:
- Limited to one copilot and one agent; findings may not generalize to other tools.
- Participants were novice agent users; results may differ for experienced users.
- Study duration and task scope were constrained, not fully reflecting real-world workflows.
- Future work should explore agent transparency, balanced proactivity, and multi-tasking workflows.
Summary
This study provides the first controlled comparison of AI copilots (GitHub Copilot) and coding agents (OpenHands) in realistic software development tasks. Agents demonstrated significant productivity gains, reducing user effort and enabling new task completion. However, user experience challenges, such as understanding agent outputs, remain. The findings highlight the potential of agents to transform developer workflows while identifying areas for improvement, including transparency and balanced proactivity. These insights are broadly applicable to designing better human-agent interactions in software engineering and beyond.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 100%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 86%
DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
CHI '26· Human-LLM Collaboration +3
- 86%
When Help Hurts: Verification Load and Fatigue with AI Coding Assistants
CHI '26· Human-LLM Collaboration +3
- 83%
Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making
CHI '23· Human-LLM Collaboration +1
- 83%
Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts
CHI '23· Human-LLM Collaboration +1
- 83%
"Are You Really Sure?'' Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making
CHI '24· Explainable AI (XAI) +1
- 83%
Automatic Macro Mining from Interaction Traces at Scale
CHI '24· Human-LLM Collaboration +1
- 83%
Interactive Debugging and Steering of Multi-Agent AI Systems
CHI '25· Human-LLM Collaboration +2
- 83%
Which Contributions Deserve Credit? Perceptions of Attribution in Human-AI Co-Creation
CHI '25· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)