"What It Wants Me To Say": Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models
Honorable MentionAuthors
Document Title
“What It Wants Me To Say”: Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models
Document Information
- Subject Area: Human-Computer Interaction, Natural Language Programming, Applications of Large Language Models (LLMs)
- Keywords: Natural Language Programming, Tabular Data Processing, Human-Computer Interaction, Large Language Models, Natural Language Interfaces, Abstraction Matching
Research Background and Problem
- Large Language Models (LLMs) enable users to generate code through natural language, but users often struggle to effectively articulate their intentions in plain language, particularly when choosing the appropriate level of abstraction for code generation models. This challenge is referred to as the "abstraction matching" problem.
- Current natural language interfaces face challenges, such as significant discrepancies between user descriptions and model interpretations when attempting to generate code. Additionally, existing models often exhibit inconsistent understanding of natural language instructions, leading to ambiguous or erroneous code generation.
- In the context of table-driven data analysis, non-expert programming users face considerable difficulties in describing their intentions in natural language to generate Python code.
- The motivation of this research is to design a method that helps users better understand how to "communicate" effectively with the model, thereby improving user experience and productivity.
Solution
- Method: Propose a "Grounded Abstraction Matching" approach, which maps users' linguistic intentions to system actions (e.g., code) and systematically translates the generated code back into editable natural language (grounded statements). This helps users learn how to effectively express their intentions to the system.
- Innovations:
- Introduce grounded corresponding statements to help users understand the code generation process.
- Provide a systematic and predictable way to align users' natural language with the abstraction level of the model's operational capabilities.
- Support error recovery and result verification, enhancing users' trust and control over the system.
- Implementation Steps:
- Design and implement two systems: one supporting abstraction matching with grounded statements and another based on non-grounded methods recommended in existing literature.
- Develop a data analysis system that supports natural language operations on tabular data, using the Codex code generation model to provide solutions.
- Employ a systematic algorithm to convert generated code into grounded statements.
- Collaborate with users iteratively and conduct experiments to compare the effectiveness of the two approaches.
Research Outcomes
- Specific Outcomes:
- Proposed an innovative "Grounded Abstraction Matching" method and applied it to real-world table analysis tasks in the context of natural language programming for data analysis.
- User studies demonstrated that compared to non-grounded methods, the grounded approach helps users better understand the system's operational logic and improves error recovery capabilities.
- Designed three real-world tasks to measure the success rate, error patterns, and user perceptions of different methods.
- Comparative Advantages:
- Users employing the grounded method were better able to interpret feedback and adjust their natural language inputs to generate correct code.
- Users exhibited improved capabilities, reflected in stronger trust and the ability to learn efficient linguistic expressions to describe task intentions.
- Experimental or Evaluation Results:
- The grounded system outperformed the non-grounded system in task completion, especially when users needed to identify partially correct outputs to rectify errors.
- Based on NASA TLX and System Usability Scale metrics, users reported moderate cognitive load and high acceptance of the system.
- In experimental tasks, the grounded system effectively reduced certain error patterns, such as input-output mismatches and logical selection biases.
- Limitations and Future Directions:
- The current implementation supports a limited range of code operations, primarily covering parts of the Pandas library. Future work could explore support for more programming language APIs.
- Experiments were conducted in a controlled lab environment; longitudinal studies and large-scale real-world applications are needed for further validation.
- Optimization of grounded statements remains an area for improvement, such as leveraging LLMs to generate more precise explanatory statements.
Conclusion
This paper successfully reduces barriers for non-expert users in natural language programming by applying "Grounded Abstraction Matching," expanding the boundaries of generative AI and user interaction design. Future research will further explore how to implement this approach in more complex scenarios, adapting to different language models and diverse user groups.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do users effectively communicate with code-generating LLMs, especially when choosing appropriate abstraction levels?Category: LLM Code Generation and Programming AssistantsSimilar questionsarrow_forward
- Can introducing embodied abstraction matching help users better understand code generation and improve error recovery?Category: LLM Code Generation and Programming AssistantsSimilar questionsarrow_forward
- In natural language programming for tabular data analysis, how can systems be designed to support users' intent expression?Category: LLM Code Generation and Programming AssistantsSimilar questionsarrow_forward
Practical Problems
1- Non-professional programming users struggle to generate accurate code through natural language.Category: LLM Code Generation and Programming AssistantsSimilar questionsarrow_forward
- 100%
Discovering the Syntax and Strategies of Natural Language Programming with Generative Language Models
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 100%
D-Twins: Your Digital Twin Designed for Real-Time Boredom Intervention
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 100%
Take It, Leave It, or Fix It: Measuring Productivity and Trust in Human-AI Collaboration
IUI '24· Generative AI (Text, Image, Music, Video) +1
- 80%
Method for Exploring Generative Adversarial Networks (GANs) via Automatically Generated Image Galleries
CHI '21· Generative AI (Text, Image, Music, Video) +2
- 80%
MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 80%
Ivie: Lightweight Anchored Explanations of Just-Generated Code
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 80%
Generative AI Uses and Risks for Knowledge Workers in a Science Organization
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 80%
How the Role of Generative AI Shapes Perceptions of Value in Human-AI Collaborative Work
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 80%
Exploring Empty Spaces: Human-in-the-Loop Data Augmentation
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 80%
GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
CHI '25· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)