Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming
Honorable MentionAuthors
Title of the Paper
Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming
Paper Information
- Subject Area: Human-Computer Interaction, AI-Assisted Programming
- Keywords: AI-assisted programming, Copilot, user state modeling, programming behavior analysis, programming efficiency, human-AI collaboration, code generation models, testing and evaluation, interface design, telemetry data
Research Background and Problem Statement
-
Problems and Challenges:
- Code recommendation systems (e.g., Copilot and CodeWhisperer) can enhance programmer productivity, but there is still a need to understand how users interact with these systems and how to optimize them.
- Current metrics (e.g., suggestion acceptance rate, reduction in characters typed) fail to fully capture the complexity of user interactions with recommendation systems.
- There is a lack of tools for detailed classification of programmer activities, making it difficult to uncover potential efficiency losses and time costs.
-
Necessity:
- Optimizing code recommendation systems is not only about efficiency but also about improving user experience and productivity.
- A deeper understanding of user behavior can help design more efficient interfaces and models to better serve developers.
-
Research Motivation and Related Work:
- The potential of AI-powered code generation models (e.g., GPT and Codex) suggests that these systems may transform software development practices.
- Previous studies have shown that programmers perceive productivity improvements but also highlight the need for more granular data beyond task completion time evaluations.
- This study systematically analyzes interaction behaviors with CodeRec systems (including Copilot) and proposes a new activity classification framework, CUPS (CodeRec User Programming States).
Solution
-
Methodology:
- The authors propose the CUPS (CodeRec User Programming States) classification framework, which categorizes programmer behaviors when interacting with code recommendation systems into 12 states.
- Through experimental design, relevant telemetry data, screen recordings, and user-provided labels were collected to analyze programming patterns.
- Design optimization suggestions are provided to reduce inefficiencies in interactions.
-
Innovations:
- A novel, multi-level user behavior classification framework (CUPS) is proposed, capable of capturing fine-grained interaction behaviors while providing meaningful summaries of overall user activities.
- Combines user-provided labels and telemetry data to analyze behaviors in programming environments.
- Visualizes user behavior timelines and state transitions, offering intuitive insights into efficiency losses.
-
Implementation Steps and Techniques:
- Develop the CUPS labeling tool, enabling users to review programming sessions and annotate specific states.
- Design experiments to collect data from 21 programmers completing tasks in a Copilot environment, including video playback, self-labeling, and telemetry records.
- Analyze the data to generate state-classified timelines and state transition diagrams, revealing behavioral patterns.
- Propose interface design and metric optimization recommendations based on observations, such as introducing state prediction and personalized interaction optimization.
Research Findings
-
Specific Findings:
- Introduced the CUPS classification framework, comprising 12 user states (e.g., "verifying suggested code," "writing new functionality").
- Analyzed time distribution across states from 3,137 labeled samples, uncovering behavioral patterns during Copilot usage.
- Found that Copilot-related tasks (e.g., verifying suggestions, handling delays) accounted for more than half of programming time.
- Provided design recommendations, such as reducing suggestion generation time and improving interfaces to adapt to users' current states.
-
Comparison with Existing Solutions and Advantages:
- Compared to traditional task completion time measurements, this study reveals specific interaction costs, such as the time required to verify and edit suggested code.
- The authors introduced detailed categorizations of behaviors related to Copilot, enabling more precise identification of inefficiency points.
- Proposed new metrics, such as adjusted acceptance rate and verification time, including the time spent verifying and editing after accepting suggestions.
-
Experimental and Evaluation Results:
- The average time spent verifying Copilot's suggested code was significantly longer (approximately five times) than simply observing the suggestions.
- Programmers spent more time in states related to interactions with the recommendation system (51.5% of programming time).
- User behavior patterns indicated that when suggestion quality was insufficient, delays in programming tasks significantly increased.
-
Limitations and Future Directions:
- Limitations include the constrained experimental task scenarios, lack of coverage for all programming languages, and the absence of long-term interaction costs (e.g., code security issues).
- Future work may include developing more sophisticated state prediction models, studying the impact of different Copilot versions, applying the CUPS method to other AI-assisted tools (e.g., writing or legal assistants), and evaluating long-term productivity impacts.
Conclusion
This study systematically analyzed programmers' interaction behaviors with code recommendation systems, proposed the CUPS framework for classifying interaction behaviors, and used empirical data to generate detailed analyses for optimizing the design and metrics of AI-assisted programming.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can users' programming states in code recommendation systems (such as Copilot) be systematically classified and modeled?Category: Coding Assistants and Multi-Turn Code SupportSimilar questionsarrow_forward
- What specific interaction efficiency losses do users experience when using code recommendation systems?Category: Coding Assistants and Multi-Turn Code SupportSimilar questionsarrow_forward
- Can interface design and interaction methods of code recommendation systems be improved to optimize user experience and productivity?Category: Coding Assistants and Multi-Turn Code SupportSimilar questionsarrow_forward
Practical Problems
1- Programmers spend substantial time verifying and editing suggested code when using code recommendation tools.Category: Coding Assistants and Multi-Turn Code SupportSimilar questionsarrow_forward
- 80%
Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication
CHI '23· Human-LLM Collaboration +1
- 80%
VAL: Interactive Task Learning with GPT Dialog Parsing
CHI '24· Human-LLM Collaboration +1
- 80%
Need Help? Designing Proactive AI Assistants for Programming
CHI '25· Human-LLM Collaboration +1
- 71%
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
IUI '26· Human-LLM Collaboration +3
- 67%
Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making
CHI '23· Human-LLM Collaboration +1
- 67%
Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts
CHI '23· Human-LLM Collaboration +1
- 67%
Automatic Macro Mining from Interaction Traces at Scale
CHI '24· Human-LLM Collaboration +1
- 67%
"If the Machine Is As Good As Me, Then What Use Am I?" – How the Use of ChatGPT Changes Young Professionals' Perception of Productivity and Accomplishment
CHI '24· Human-LLM Collaboration +1
- 67%
Interactive Debugging and Steering of Multi-Agent AI Systems
CHI '25· Human-LLM Collaboration +2
- 67%
GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
CHI '25· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)