AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools

Honorable Mention
Generative AI (Text, Image, Music, Video)Human-LLM CollaborationUI/UX DesignersAI/ML Researchers & Engineers

Research Background and Issues

  • Identified Problems and Challenges: The authors highlight several core challenges with existing chat-based generative AI interfaces: boundary-based dialogues limit users' ability to express and refine ambiguous intentions, making it difficult to explore multiple options or iterate quickly. Users struggle to clearly articulate complex goals and must rely on trial-and-error to improve workflows. This linear workflow constrains non-linear creative activities, and users find it hard to understand AI's potential and the possible outcomes it can generate.
  • Significance of the Problem: As generative AI becomes increasingly prevalent in creative design, enhancing users' ability to express intentions, fostering diverse outcomes, and enabling non-linear workflows have become critical needs to address user pain points.
  • Research Motivation and Related Work: Previous studies have attempted to address some of these issues through multimodal interactions and enriched user interfaces, but they have not fully met the demands for complex behaviors such as intention expression, divergence, guidance, and exploration. Drawing on interaction design theories (e.g., tool interaction models), the authors propose a novel approach tailored to generative AI application scenarios.

Solution

  • Proposed Solution: The authors introduce AI-Instruments, which treat "prompts" as interface objects based on three core principles: Reification, Reflection, and Grounding through examples. These tools are designed as graphical interactive objects to support non-linear and direct manipulation workflows.
  • Innovative Features:
    • Reification: Transforming user intentions into interactive graphical interface objects, enabling reuse and direct manipulation.
    • Reflection Principle: Allowing users to explore not only multiple possible outcomes (reflection on responses) but also different ways of expressing intentions (reflection on intentions).
    • Grounding Principle: Enabling users to extract and apply specific attributes, such as visual styles or content details, from example content or other tools.
  • Implementation Steps and Key Technologies:
    • Tool Instances: Four technical probes are provided: Fragments, Generative Containers, Transformative Lenses, and Fillable Brushes.
    • Technical Implementation: Leveraging generative language and visual models (e.g., GPT-4o and Stable Diffusion) to generate and combine various interaction methods.
    • Workflow Design: Using the tools to accomplish tasks such as exploration (e.g., diverse generation), fine-tuning (e.g., local image modifications), and complex content synthesis.

Research Outcomes

  • Specific Results:
    • Functionality Demonstration of Technical Probes: The authors developed four instances of AI-Instruments, showcasing how the new model supports multi-tasking, such as content composition, iterative modification, style transfer, and content expansion.
    • User Feedback: In user studies, 12 participants generally agreed that AI-Instruments helped address issues like ambiguous intention expression, exploration, and result guidance. Compared to traditional chat interfaces, these tools better met the needs of complex creative work.
  • Advantages Over Existing Solutions: AI-Instruments provide a more flexible and user-friendly interface, supporting non-linear and direct manipulation, enhancing users' exploratory capabilities, iteration efficiency, and content generation quality. For instance, Generative Containers and Fragments help users develop multiple creative directions, while Fillable Brushes enable precise content adjustments.
  • Limitations and Future Directions:
    • Limitations:
      • Some users reported that GUIs could be more complex to operate than text-based prompts, leading to slightly lower efficiency for simple tasks.
      • The tools' generation capabilities might result in overly complex interfaces, necessitating further optimization of tool quantity and hierarchy.
    • Future Directions:
      • Expanding AI-Instruments to multimodal content (e.g., text, audio, or video) and enhancing their ability to handle complex projects (e.g., multi-content synthesis).
      • Further exploring the integration of tools with chat-based interactions to provide more flexible workflows for precise control and simple operations.
      • Improving version control and history tracking to address traceability issues caused by the non-deterministic outputs of current generative tools.

Through this research, the authors provide theoretical support and practical tools for designing next-generation generative AI interaction interfaces, introducing new possibilities for users in creative content generation and iterative work.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188523/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714259
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Honorable Mention
group
Authors
9 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration
work
Professions
UI/UX Designers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers