AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools
Honorable MentionAuthors
Research Background and Issues
- Identified Problems and Challenges: The authors highlight several core challenges with existing chat-based generative AI interfaces: boundary-based dialogues limit users' ability to express and refine ambiguous intentions, making it difficult to explore multiple options or iterate quickly. Users struggle to clearly articulate complex goals and must rely on trial-and-error to improve workflows. This linear workflow constrains non-linear creative activities, and users find it hard to understand AI's potential and the possible outcomes it can generate.
- Significance of the Problem: As generative AI becomes increasingly prevalent in creative design, enhancing users' ability to express intentions, fostering diverse outcomes, and enabling non-linear workflows have become critical needs to address user pain points.
- Research Motivation and Related Work: Previous studies have attempted to address some of these issues through multimodal interactions and enriched user interfaces, but they have not fully met the demands for complex behaviors such as intention expression, divergence, guidance, and exploration. Drawing on interaction design theories (e.g., tool interaction models), the authors propose a novel approach tailored to generative AI application scenarios.
Solution
- Proposed Solution: The authors introduce AI-Instruments, which treat "prompts" as interface objects based on three core principles: Reification, Reflection, and Grounding through examples. These tools are designed as graphical interactive objects to support non-linear and direct manipulation workflows.
- Innovative Features:
- Reification: Transforming user intentions into interactive graphical interface objects, enabling reuse and direct manipulation.
- Reflection Principle: Allowing users to explore not only multiple possible outcomes (reflection on responses) but also different ways of expressing intentions (reflection on intentions).
- Grounding Principle: Enabling users to extract and apply specific attributes, such as visual styles or content details, from example content or other tools.
- Implementation Steps and Key Technologies:
- Tool Instances: Four technical probes are provided: Fragments, Generative Containers, Transformative Lenses, and Fillable Brushes.
- Technical Implementation: Leveraging generative language and visual models (e.g., GPT-4o and Stable Diffusion) to generate and combine various interaction methods.
- Workflow Design: Using the tools to accomplish tasks such as exploration (e.g., diverse generation), fine-tuning (e.g., local image modifications), and complex content synthesis.
Research Outcomes
- Specific Results:
- Functionality Demonstration of Technical Probes: The authors developed four instances of AI-Instruments, showcasing how the new model supports multi-tasking, such as content composition, iterative modification, style transfer, and content expansion.
- User Feedback: In user studies, 12 participants generally agreed that AI-Instruments helped address issues like ambiguous intention expression, exploration, and result guidance. Compared to traditional chat interfaces, these tools better met the needs of complex creative work.
- Advantages Over Existing Solutions: AI-Instruments provide a more flexible and user-friendly interface, supporting non-linear and direct manipulation, enhancing users' exploratory capabilities, iteration efficiency, and content generation quality. For instance, Generative Containers and Fragments help users develop multiple creative directions, while Fillable Brushes enable precise content adjustments.
- Limitations and Future Directions:
- Limitations:
- Some users reported that GUIs could be more complex to operate than text-based prompts, leading to slightly lower efficiency for simple tasks.
- The tools' generation capabilities might result in overly complex interfaces, necessitating further optimization of tool quantity and hierarchy.
- Future Directions:
- Expanding AI-Instruments to multimodal content (e.g., text, audio, or video) and enhancing their ability to handle complex projects (e.g., multi-content synthesis).
- Further exploring the integration of tools with chat-based interactions to provide more flexible workflows for precise control and simple operations.
- Improving version control and history tracking to address traceability issues caused by the non-deterministic outputs of current generative tools.
- Limitations:
Through this research, the authors provide theoretical support and practical tools for designing next-generation generative AI interaction interfaces, introducing new possibilities for users in creative content generation and iterative work.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can GenAI interaction be designed for users to efficiently express vague intent and iterate quickly?Category: Creative Workflows and Multi-Stage PipelinesSimilar questionsarrow_forward
- How can GenAI interfaces support nonlinear creative workflows and diverse exploration?Category: Creative Workflows and Multi-Stage PipelinesSimilar questionsarrow_forward
- How can graphical tools enhance users' understanding of GenAI potential?Category: Creative Workflows and Multi-Stage PipelinesSimilar questionsarrow_forward
Practical Problems
1- Creators struggle to clearly express intent and quickly explore diverse AI-generated results.Category: Creative Workflows and Multi-Stage PipelinesSimilar questionsarrow_forward
- 100%
IntentTuner: An Interactive Framework for Integrating Human Intentions in Fine-tuning Text-to-Image Generative Models
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 80%
Is It AI or Is It Me? Understanding Users’ Prompt Journey with Text-to-Image Generative AI Tools
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 80%
MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 80%
How the Role of Generative AI Shapes Perceptions of Value in Human-AI Collaborative Work
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 80%
GANzilla: User-Driven Direction Discovery in Generative Adversarial Networks
UIST '22· Generative AI (Text, Image, Music, Video) +1
- 75%
User Experience Design Professionals’ Perceptions of Generative Artificial Intelligence
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 75%
Cells, Generators, and Lenses: Design Framework for Object-Oriented Interaction with Large Language Models
UIST '23· Human-LLM Collaboration
- 75%
Patchview: LLM-powered Worldbuilding with Generative Dust and Magnet Visualization
UIST '24· Generative AI (Text, Image, Music, Video) +1
- 67%
ReadingQuizMaker: A Human-NLP Collaborative System to Support Instructors Design High Quality Reading Quiz Questions
CHI '23· Generative AI (Text, Image, Music, Video) +2
- 67%
ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
CHI '24· Voice User Interface (VUI) Design +2
Based on Jaccard similarity of research subtopics & professions (≥60%)