ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts
Authors
Large language models (LLMs) enable the rapid generation of data wrangling scripts based on natural language instructions, but these scripts may not fully adhere to user-specified requirements, necessitating careful inspection and iterative refinement. Existing approaches primarily assist users in understanding script logic and spotting potential issues themselves, rather than providing direct validation of correctness. To enhance debugging efficiency and optimize the user experience, we develop ViseGPT, a tool that automatically extracts constraints from user prompts to generate comprehensive test cases for verifying script reliability. The test results are then transformed into a tailored Gantt chart, allowing users to intuitively assess alignment with semantic requirements and iteratively refine their scripts. Our design decisions are informed by a formative study (N=8) that explores user practices and challenges. We further evaluate the effectiveness and usability of ViseGPT through a user study (N=18). Results indicate that ViseGPT significantly improves debugging efficiency for LLM-generated data-wrangling scripts, enhances users’ ability to detect and correct issues, and streamlines the workflow experience.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Lexara: A User-Centered Toolkit for Evaluating Large Language Models for Conversational Visual Analytics
CHI '26· Human-LLM Collaboration +2
- 67%
A Human-Computer Collaborative Editing Tool for Conceptual Diagrams
CHI '23· Human-LLM Collaboration +1
- 67%
Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
CHI '25· Human-LLM Collaboration +1
- 67%
Wikipedia ORES Explorer: Visualizing Trade-offs For Designing Applications With Machine Learning API
DIS '21· Explainable AI (XAI) +1
- 67%
Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking
IUI '24· Human-LLM Collaboration +1
- 60%
MAPLE: Mobile App Prediction Leveraging Large Language Model Embeddings
UbiComp '24· Human-LLM Collaboration
- 60%
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation
UIST '24· Human-LLM Collaboration
Based on Jaccard similarity of research subtopics & professions (≥60%)