Investigating Explainability of Generative Models for Code through Scenario-based Design
Authors
Title of the Paper
Investigating Explainability of Generative AI for Code through Scenario-Based Design
Paper Information
- Domain: Explainability design for generative AI, specifically in the context of code generation
- Keywords: Generative AI, code generation tools, explainable AI, human-centered AI, scenario-based design
Research Background and Problem
-
Problems and Challenges:
- Unlike traditional discriminative machine learning models, generative AI (GenAI) focuses on generating artifacts (e.g., code snippets) rather than making decisions, which introduces unique explainability requirements.
- While generative code models like GitHub CoPilot and TransCoder show potential, there is still limited understanding of how users comprehend and effectively utilize these tools.
- Existing research on explainability in generative models is severely lacking, with most studies focused on discriminative models, neglecting real-world user scenarios and needs.
-
Significance:
- Generative AI is rapidly advancing the automation of software development in tasks like code auto-completion, natural language-to-code translation, and code translation.
- Understanding users' explainability needs when using generative AI for code can help develop more user-friendly and efficient tools, enhancing productivity and reducing errors in software development.
-
Research Motivation and Related Work:
- The motivation stems from "human-centered explainable AI" (Human-Centered XAI), which emphasizes understanding user needs in specific scenarios and designing tools driven by those needs.
- Existing frameworks, such as the problem-driven XAI approach proposed by Liao et al., provide theoretical foundations for investigating user needs but have not been deeply explored in the context of generative AI for code generation.
Solution
-
Methods/Solutions:
- Employ scenario-based design (SBD) combined with problem-driven XAI design to explore explainability needs in generative AI for code scenarios.
- Focus on three typical use cases: code translation, auto-completion, and natural language-to-code generation, collecting feedback from 43 software engineers through nine participatory workshops.
- Propose four categories of XAI feature designs: AI documentation support, model uncertainty indication, attention distribution visualization, and social transparency, with iterative design recommendations.
-
Innovations:
- Identified 11 categories of explainability needs for generative AI, including new requirements specific to code generation (e.g., input scope, control options, system limitations).
- Provided concrete recommendations to enhance the explainability design of generative code models, such as documentation support and phased perspectives on social transparency.
-
Implementation Steps/Key Techniques:
- Design low-fidelity prototypes with real-world examples of generative AI outputs for participants.
- Collect users' XAI needs using a problem-driven design approach, organizing needs through an "aggregate-classify-vote" process.
- Discuss four major XAI feature designs during workshops, recording user feedback and suggestions for improvement.
Research Findings
-
Specific Findings:
- Identified 11 categories of explainability needs for generative AI, surpassing prior research concepts (e.g., data, performance) and introducing new requirements specific to code generation (e.g., system requirements, control over code characteristics).
- Proposed AI documentation guidelines tailored to real-world software engineering needs, including performance metrics, supported languages and frameworks, and training data sources.
- Validated the potential value of localized explanations based on attention distribution and uncertainty indication in generative code AI.
-
Comparative Advantages:
- Unlike existing XAI research focused on decision support, this study provides a detailed analysis of the needs for generative AI in complex input-output spaces.
- Offers refined feature suggestions (e.g., displaying social transparency at different stages of the code lifecycle) to better support user tasks.
-
Experimental or Evaluation Results:
- Engineers widely acknowledged the value of AI uncertainty indication and social transparency but emphasized the need for more interaction, such as displaying alternative AI-generated code options or the sources of uncertainty.
- Data needs for documentation support suggest a focus on example-based teaching and contextual explanations to address knowledge gaps among non-technical users.
-
Limitations and Future Directions:
- Generative AI models still exhibit high error rates, requiring users to perform extensive post-review and revisions to mitigate risks.
- Future research should explore the usability of generative AI in scenarios involving non-technical users and develop interaction-centered explanation methods.
- Continuously improve social transparency design from a dynamic organizational perspective, integrating it into different stages of the code development lifecycle.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What explainability requirements exist for GenAI in code generation scenarios?Category: Code and Programming-Assisted CreationSimilar questionsarrow_forward
- How can GenAI for code generation improve usability and user understanding through scenario-based design?Category: Code and Programming-Assisted CreationSimilar questionsarrow_forward
- Which explainability features of generative code models can improve software development efficiency and reduce errors?Category: Code and Programming-Assisted CreationSimilar questionsarrow_forward
Practical Problems
1- Programmers struggle to understand model outputs and their limitations when using GenAI to write code.Category: Code and Programming-Assisted CreationSimilar questionsarrow_forward
- 100%
Ivie: Lightweight Anchored Explanations of Just-Generated Code
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 83%
Exploring the Innovation Opportunities for Pre-trained Models
DIS '25· Generative AI (Text, Image, Music, Video) +2
- 80%
Discovering the Syntax and Strategies of Natural Language Programming with Generative Language Models
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 80%
"What It Wants Me To Say": Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models
CHI '23· Generative AI (Text, Image, Music, Video) +1
- 80%
Validating AI-Generated Code with Live Programming
CHI '24· Human-LLM Collaboration +1
- 80%
Exploring the Design Space of Real-time LLM Knowledge Support Systems: A Case Study of Jargon Explanations
CHI '25· Human-LLM Collaboration +1
- 80%
D-Twins: Your Digital Twin Designed for Real-Time Boredom Intervention
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 80%
From Discovery to Adoption: Understanding the ML Practitioners' Interpretability Journey
DIS '23· Generative AI (Text, Image, Music, Video) +1
- 80%
Take It, Leave It, or Fix It: Measuring Productivity and Trust in Human-AI Collaboration
IUI '24· Generative AI (Text, Image, Music, Video) +1
- 80%
Less or More: Towards Glanceable Explanations for LLM Recommendations Using Ultra-Small Devices
IUI '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)