New Enactions of Expertise: Software Engineers’ Evaluation and Demonstration of Coding Expertise with AI Coding Assistants
Authors
Paper Title
Evolving Enactions of Expertise: Software Engineers’ Evaluation and Demonstration of Coding Expertise with AI Coding Assistants
Publication Info
- Topic area: The impact of AI coding assistants on the evaluation and demonstration of coding expertise in software engineering.
- Keywords: AI coding assistants, coding expertise, live coding interviews, software engineering, evaluation criteria, productivity, enaction of expertise, task comprehension, code implementation, AI-human collaboration.
Background and Problem
- Problem / challenge: The integration of AI coding assistants into software engineering workflows raises questions about how coding expertise is demonstrated and evaluated, as traditional workflows and evidence of expertise are altered or obscured.
- Significance: Understanding these changes is crucial for aligning hiring and training practices with evolving industry norms and ensuring fair and effective evaluation of coding expertise.
- Motivation and related work: Prior research has focused on the capabilities of AI coding assistants and their impact on productivity, but little is known about how they reshape the demonstration and evaluation of coding expertise. This study builds on the concept of expertise as socially and interactionally constituted, exploring how it evolves with AI tools.
Solution
- Proposed approach: The study investigates how software engineers demonstrate and evaluate coding expertise in AI-assisted contexts through simulated live coding interviews, where evaluators assess candidates using AI tools.
- Novelty:
- Empirical insights into how coding expertise is demonstrated and evaluated with AI assistants.
- Introduction of the concept of "enaction of expertise" to capture evolving practices in AI-assisted coding.
- Design and organizational recommendations for aligning evaluation and training with AI-driven changes in coding expertise.
- Procedure and key techniques:
- Conducted 12 simulated live coding interviews with 16 participants, pairing evaluators and candidates.
- Evaluators designed tasks and criteria, then revised them to account for AI tool availability.
- Candidates used various AI tools during the interviews, with their interactions and outputs analyzed.
- Data included pre- and post-interview reflections, live coding session recordings, and thematic analysis of interactions.
Results
- Concrete findings:
- Evaluators used familiar criteria (e.g., task comprehension, code explanation, error handling) but required additional evidence when AI tools were involved.
- New enactions of expertise emerged, such as tool selection, prompting practices, and use of AI-generated outputs.
- Established workflows like incremental code development and error handling were diminished, reducing opportunities to demonstrate expertise.
- Higher productivity expectations arose with AI tools, often conflicting with traditional indicators of expertise.
- Advantage over baselines: The study highlights nuanced shifts in coding practices and evaluation criteria that are not captured by traditional coding interviews or productivity-focused studies.
- Experiments / evaluation:
- Participants included 16 software engineers with varying levels of experience and AI tool familiarity.
- Tasks spanned backend development, data science, and AI integration, with candidates using tools like ChatGPT, Copilot, and Perplexity.
- Metrics included task completion, interaction logs, evaluator criteria, and qualitative reflections.
- Limitations and future work:
- Evaluators’ familiarity with candidates’ chosen AI tools may have influenced assessments.
- Variability in task design and complexity could affect findings.
- The study's live-coding format does not fully capture the collaborative and iterative nature of real-world software engineering.
- Future research should explore longitudinal and team-based workflows, as well as standardized tasks.
Summary
This study examines how AI coding assistants reshape the evaluation and demonstration of coding expertise in software engineering. While evaluators continued to use familiar criteria, they required additional evidence due to new enactions of expertise, such as tool selection and prompting practices, and diminished visibility of traditional workflows like error handling. The study also found heightened productivity expectations, often misaligned with expertise indicators. These findings suggest the need for tools and training to support both candidates and evaluators in navigating evolving coding practices. By introducing the concept of "enaction of expertise," the study provides a framework for understanding and adapting to these shifts in an AI-driven coding landscape.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 100%
From Junior to Senior: Allocating Agency and Navigating Professional Growth in Agentic AI-Mediated Software Engineering
CHI '26· Human-LLM Collaboration +2
- 100%
Developer Interaction Patterns with Proactive AI: A Five-Day Field Study
IUI '26· AI-Assisted Decision-Making & Automation +2
- 83%
Vibe Coding Entanglements – Repositioning Boundaries of Intention, Authorship, and Responsibility in Programming with Generative AI
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 83%
Gemini at Work: Knowledge Workers' Perceptions and Assessment of Productivity Gains
DIS '25· Generative AI (Text, Image, Music, Video) +3
- 83%
PREDILECT: Preferences Delineated with Zero-Shot Language-based Reasoning in Reinforcement Learning
HRI '24· Generative AI (Text, Image, Music, Video) +2
- 83%
GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
UIST '24· Generative AI (Text, Image, Music, Video) +2
- 80%
Discovering the Syntax and Strategies of Natural Language Programming with Generative Language Models
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 80%
Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication
CHI '23· Human-LLM Collaboration +1
- 80%
"What It Wants Me To Say": Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models
CHI '23· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)