Better Together? An Evaluation of AI-Supported Code Translation

Generative AI (Text, Image, Music, Video)Human-LLM CollaborationSoftware Engineers & Developers

Document Title

Better Together? An Evaluation of AI-Supported Code Translation

Document Information

  • Subject Area: Human-AI Collaboration, Code Translation, Design and Application of Generative AI Systems
  • Keywords: Code Translation, Human-AI Co-Creation, Generative AI, Imperfect AI, Software Engineering

Research Background and Problem

  • Identified Problems:

    1. State-of-the-art generative AI models (e.g., TransCoder and Codex) exhibit errors or incomplete outputs in code generation and translation tasks, with accuracy rates ranging from 30% to 70%.
    2. Whether software engineers can improve productivity by collaborating with imperfect generative AI models remains underexplored.
  • Significance of the Research: Leveraging generative AI to enhance code quality and engineering efficiency holds significant potential. However, it is crucial to determine how to design human-AI interactions to maximize collaborative outcomes. Additionally, understanding the processes of learning and working with flawed AI outputs has both theoretical and practical implications for designing better AI-assisted interfaces in the future.

  • Motivation and Related Work:

    1. While generative AI has achieved some success in code translation and completion, research on whether these tools genuinely enhance the efficiency and outcomes of software engineers is insufficient.
    2. This study experimentally demonstrates the usability and efficiency improvements of generative AI in code translation tasks and analyzes its side effects and user experience.

Solution

  • Proposed Approach: The authors designed a comparative experiment to investigate whether software engineers could achieve more efficient outputs in Java-to-Python code translation tasks with the support of generative AI models.

  • Innovations:

    1. Examined the impact of the quantity and quality of generative AI outputs on collaborative performance, clarifying whether multiple translation options aid user decision-making.
    2. Tested whether AI-assisted translation could shift the task focus from "code production" to "code review" under time-constrained conditions.
  • Implementation Steps and Key Techniques:

    1. Used the TransCoder model to generate Java-to-Python translations for data structures (Trie and Priority Queue), setting conditions for "poor" and "good" output quality.
    2. Conducted experiments with 32 professional software engineers, who completed tasks under two conditions (without AI support vs. with AI support) via random pairing.
    3. Evaluated the quality of output code, quantifying error rates, accuracy, and over 1,700 categorized programming errors (e.g., translation errors, language errors, documentation omissions).

Research Findings

  • Specific Results:

    1. Improved Translation Performance:
      • AI-supported code error rates significantly decreased (from 63% to 31%).
      • The proportion of correctly translated functions increased from 34.1% to 50.9%.
    2. User Behavior and Habits:
      • Users shifted tasks from "code production" to "code review," reducing the workload of mechanical coding.
      • Over half of the users reported that AI translations saved time and helped them learn new programming patterns.
    3. Impact of Translation Options:
      • Providing 5 translation versions (compared to 1 version) increased workload and cognitive pressure on users.
      • "Higher quality" translation versions directly contributed to error reduction.
  • Comparison with Existing Research: This study further supports the potential advantages of human-AI collaboration in complex tasks and challenges the notion that "AI must be perfect to be useful."

  • Limitations and Future Directions:

    1. This study primarily focused on "code translation" tasks; future research could extend to other software engineering scenarios, such as code generation and documentation generation.
    2. Interaction Design Directions: Develop smarter user interfaces, such as:
      • Highlighting differences among multiple translations.
      • Implementing features akin to "code coach" or "intelligent review assistant."
    3. Integration with XAI (Explainable AI): Incorporate explainability into generative AI outputs to help developers better understand and validate machine-generated results.

Additional Discussion

  • Societal Impact:
    • Positive: Code generation AI tools can lower the barriers to professional programming, democratizing software engineering.
    • Negative: Over-automation may replace human labor; designing human-AI collaborative systems is necessary to mitigate risks.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/79972/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3490099.3511157
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration
work
Professions
Software Engineers & Developers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers