DirectGPT: A Direct Manipulation Interface to Interact with Large Language Models
Honorable MentionAuthors
Title of the Paper
DirectGPT: A Direct Manipulation Interface to Interact with Large Language Models
Paper Information
- Subject Area: Human-Computer Interaction, Artificial Intelligence, and Interface Design for Large Language Models
- Keywords: Large Language Models, User Interface, Direct Manipulation, Prompt Engineering, Editing Tasks, Graphical User Interface, User Study, Tool Integration, Human-AI Collaboration, Reversible Operations
Research Background and Problem Statement
-
Identified Issues or Challenges: Current interaction interfaces for large language models (LLMs) primarily rely on text-based conversational prompts, which lead to several problems:
- Difficulty in Expressing Intent: Using natural language to describe goals often results in ambiguity, especially for tasks involving structured information or images.
- Unpredictable Output: LLMs may produce different or unexpected results due to slight variations in prompts.
- High Trial-and-Error Costs: Users need to spend significant time adjusting prompts to achieve satisfactory results.
- Limited Reusability of Prompts: Once a task is completed using a prompt, the result is difficult to adapt for other objectives.
- Lack of Reversible Operations: Traditional conversational interfaces struggle to support interactions like undo and redo.
-
Significance: While LLMs are being widely adopted, these issues hinder effective interaction with users and limit their potential in complex tasks. Improving interaction methods can enhance user efficiency and the quality of generated content.
-
Research Motivation and Related Work:
- Numerous studies have explored tools to improve prompt engineering and interaction experiences, but they are often limited to specific tasks and fail to address the broad application scenarios of LLMs.
- This study draws inspiration from the "direct manipulation interface" principle in human-computer interaction, aiming to apply it to LLM interaction interfaces to address the aforementioned issues.
Solution
-
Proposed Method or Solution: A direct manipulation interface—DirectGPT—is proposed. By combining physical actions with prompt engineering, it translates user operations into generated prompts, improving content controllability and interaction efficiency.
- Core principles include:
- Continuous Representation: Continuously display generated objects to support direct interaction.
- Physical Actions: Users represent objects in prompts through actions like selection and drag-and-drop.
- Quick Operations: Executed prompts are transformed into tools for subsequent reuse.
- Immediate Feedback: Visual cues clarify target objects and the effects of modifications.
- Reversible Operations: Undo and redo buttons are provided to support operational flexibility.
- Core principles include:
-
Innovative Contributions:
- Improved prompt engineering by simplifying input language through spatial selection and object referencing.
- Transformed prompts into reusable tools, reducing the complexity of repetitive operations.
- Simplified language descriptions using direct manipulation, enhancing the predictability and controllability of results.
-
Implementation Steps and Key Technologies:
- Develop a DirectGPT prototype system based on ReactJS and OpenAI API.
- Implement direct manipulation functionalities for editing text, code, and SVG images.
- Adjust internal LLM prompts to accommodate user-generated objects, such as unique markers in text or object IDs in images.
Research Outcomes
-
Specific Results:
- DirectGPT enables users to generate prompts through direct interaction, excelling in editing tasks.
- User studies show that tasks completed with DirectGPT are 50% faster than with ChatGPT, using 50% fewer prompts and reducing prompt length by 72%.
- In System Usability Scale (SUS) scores, DirectGPT significantly outperformed ChatGPT (92 vs. 53).
-
Advantages:
- Clear object selection reduces ambiguity in prompt language.
- Feedback on effects enhances user control over modification targets and scope.
- Reversible operations improve trial-and-error efficiency and error recovery capabilities.
-
Experimental or Evaluation Results:
- Users demonstrated significantly better performance in task completion time, prompt quantity, and language conciseness compared to ChatGPT.
- Subjective questionnaire results indicate that DirectGPT outperforms ChatGPT in dimensions such as object selection, content output control, and operational clarity.
-
Limitations and Future Directions:
-
Limitations:
- DirectGPT's performance depends on the capabilities of current LLMs, which may impose constraints.
- Study participants were all technically skilled users familiar with programming and LLMs, so results may not generalize to non-technical users.
- Direct manipulation may have limited utility for exploratory tasks or global modifications.
-
Future Research Directions:
- Investigate the impact of direct manipulation interfaces on learning and exploratory tasks.
- Integrate demonstrative interfaces (tools for generating example operations) and other user interface interaction modes.
- Extend DirectGPT to domains like data processing, website design, and graphical user interface construction.
- Explore the application of DirectGPT's design to other generative models, such as image generation and music creation.
-
Conclusion
DirectGPT improves the interaction experience with LLMs through the principle of direct manipulation. Its innovative design helps users express intentions more efficiently, control generated content, and seamlessly integrate into traditional software. This study provides significant theoretical and practical support for designing LLM interaction interfaces and demonstrates the potential of intuitive user interfaces in collaborative artificial intelligence.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Can direct manipulation interfaces improve interaction efficiency and output control when working with large language models (LLMs)?Category: LLM Interfaces, Prompts, and Interaction UnderstandingSimilar questionsarrow_forward
- Can simplified input language and direct manipulation improve predictability and controllability of generated content?Category: LLM Interfaces, Prompts, and Interaction UnderstandingSimilar questionsarrow_forward
- How can direct manipulation interfaces support undo and redo to optimize users' trial-and-error costs?Category: LLM Interfaces, Prompts, and Interaction UnderstandingSimilar questionsarrow_forward
Practical Problems
1- Users struggle to express intent accurately in text and face high costs when adjusting language model outputs.Category: LLM Interfaces, Prompts, and Interaction UnderstandingSimilar questionsarrow_forward
- 75%
Adapting User Interfaces with Model-based Reinforcement Learning
CHI '21· Human-LLM Collaboration +2
- 75%
Selenite: Scaffolding Online Sensemaking with Comprehensive Overviews Elicited from Large Language Models
CHI '24· Human-LLM Collaboration +2
- 75%
ChainBuddy: An AI-assisted Agent System for Generating LLM Pipelines
CHI '25· Human-LLM Collaboration +1
- 75%
Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
CHI '26· Human-LLM Collaboration +2
- 75%
DataSpeck: An AI-Driven Human-in-the-Loop System for Automating Transformations in Data Conversion Workflows
CHI '26· Human-LLM Collaboration +2
- 75%
CoAutoML: User Interface Framework for Machine Learning Novices using LLM-based AutoML and Test-Driven Machine Teaching
IUI '26· AutoML Interfaces +2
- 75%
RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented Generation
UIST '25· Human-LLM Collaboration +2
- 67%
Solving Separation-of-Concerns Problems in Collaborative Design of Human-AI Systems through Leaky Abstractions
CHI '22· Human-LLM Collaboration +3
- 67%
VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
CHI '25· Human-LLM Collaboration +2
- 67%
When Help Hurts: Verification Load and Fatigue with AI Coding Assistants
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)