Better Together? An Evaluation of AI-Supported Code Translation
Authors
Document Title
Better Together? An Evaluation of AI-Supported Code Translation
Document Information
- Subject Area: Human-AI Collaboration, Code Translation, Design and Application of Generative AI Systems
- Keywords: Code Translation, Human-AI Co-Creation, Generative AI, Imperfect AI, Software Engineering
Research Background and Problem
-
Identified Problems:
- State-of-the-art generative AI models (e.g., TransCoder and Codex) exhibit errors or incomplete outputs in code generation and translation tasks, with accuracy rates ranging from 30% to 70%.
- Whether software engineers can improve productivity by collaborating with imperfect generative AI models remains underexplored.
-
Significance of the Research: Leveraging generative AI to enhance code quality and engineering efficiency holds significant potential. However, it is crucial to determine how to design human-AI interactions to maximize collaborative outcomes. Additionally, understanding the processes of learning and working with flawed AI outputs has both theoretical and practical implications for designing better AI-assisted interfaces in the future.
-
Motivation and Related Work:
- While generative AI has achieved some success in code translation and completion, research on whether these tools genuinely enhance the efficiency and outcomes of software engineers is insufficient.
- This study experimentally demonstrates the usability and efficiency improvements of generative AI in code translation tasks and analyzes its side effects and user experience.
Solution
-
Proposed Approach: The authors designed a comparative experiment to investigate whether software engineers could achieve more efficient outputs in Java-to-Python code translation tasks with the support of generative AI models.
-
Innovations:
- Examined the impact of the quantity and quality of generative AI outputs on collaborative performance, clarifying whether multiple translation options aid user decision-making.
- Tested whether AI-assisted translation could shift the task focus from "code production" to "code review" under time-constrained conditions.
-
Implementation Steps and Key Techniques:
- Used the TransCoder model to generate Java-to-Python translations for data structures (Trie and Priority Queue), setting conditions for "poor" and "good" output quality.
- Conducted experiments with 32 professional software engineers, who completed tasks under two conditions (without AI support vs. with AI support) via random pairing.
- Evaluated the quality of output code, quantifying error rates, accuracy, and over 1,700 categorized programming errors (e.g., translation errors, language errors, documentation omissions).
Research Findings
-
Specific Results:
- Improved Translation Performance:
- AI-supported code error rates significantly decreased (from 63% to 31%).
- The proportion of correctly translated functions increased from 34.1% to 50.9%.
- User Behavior and Habits:
- Users shifted tasks from "code production" to "code review," reducing the workload of mechanical coding.
- Over half of the users reported that AI translations saved time and helped them learn new programming patterns.
- Impact of Translation Options:
- Providing 5 translation versions (compared to 1 version) increased workload and cognitive pressure on users.
- "Higher quality" translation versions directly contributed to error reduction.
- Improved Translation Performance:
-
Comparison with Existing Research: This study further supports the potential advantages of human-AI collaboration in complex tasks and challenges the notion that "AI must be perfect to be useful."
-
Limitations and Future Directions:
- This study primarily focused on "code translation" tasks; future research could extend to other software engineering scenarios, such as code generation and documentation generation.
- Interaction Design Directions: Develop smarter user interfaces, such as:
- Highlighting differences among multiple translations.
- Implementing features akin to "code coach" or "intelligent review assistant."
- Integration with XAI (Explainable AI): Incorporate explainability into generative AI outputs to help developers better understand and validate machine-generated results.
Additional Discussion
- Societal Impact:
- Positive: Code generation AI tools can lower the barriers to professional programming, democratizing software engineering.
- Negative: Over-automation may replace human labor; designing human-AI collaborative systems is necessary to mitigate risks.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Can software engineers improve efficiency and output quality using GenAI in code translation tasks?Category: AI Code Translation and Error Repair SupportSimilar questionsarrow_forward
- How do the quantity and quality of GenAI outputs affect human-AI collaboration outcomes?Category: AI Code Translation and Error Repair SupportSimilar questionsarrow_forward
- Under time constraints, can AI-assisted translation shift task focus from "code production" to "code review"?Category: AI Code Translation and Error Repair SupportSimilar questionsarrow_forward
Practical Problems
1- Software engineers struggle to maintain work efficiency and output quality when facing errors in AI-generated code translations.Category: AI Code Translation and Error Repair SupportSimilar questionsarrow_forward
- 75%
Discovering the Syntax and Strategies of Natural Language Programming with Generative Language Models
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 75%
Slide4N: Creating Presentation Slides from Computational Notebooks with Human-AI Collaboration
CHI '23· Generative AI (Text, Image, Music, Video) +1
- 75%
"What It Wants Me To Say": Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models
CHI '23· Generative AI (Text, Image, Music, Video) +1
- 75%
Generative and Malleable User Interfaces with Generative and Evolving Task-Driven Data Model
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 75%
The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 75%
GeneyMAP: Exploring the Potential of GenAI to Facilitate Mapping User Journeys for UX Design
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 75%
D-Twins: Your Digital Twin Designed for Real-Time Boredom Intervention
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 75%
Cracking the code - Co-coding with AI in creative programming education
C&C '22· Generative AI (Text, Image, Music, Video) +1
- 75%
Take It, Leave It, or Fix It: Measuring Productivity and Trust in Human-AI Collaboration
IUI '24· Generative AI (Text, Image, Music, Video) +1
- 67%
PANDALens: Towards AI-Assisted In-Context Writing on OHMD During Travels
CHI '24· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)