Compositional Structures as Substrates for Human-AI Co-creation Environment: A Design Approach and A Case Study

Generative AI (Text, Image, Music, Video)Creative Collaboration & Feedback SystemsMusicians, DJs & Sound DesignersFilm & Animation ProducersUI/UX Designers

Research Background and Issues

  • Issues and Challenges:

    • Current human-AI co-creation environments often rely on a linear "prompt-result" interaction model, lacking support for exploration, planning, iteration, and control or inspection of AI-generated content during the creative process.
    • In complex content creation processes, such as video production, effectively organizing and managing multidimensional information (e.g., narrative, temporal, spatial classification) remains a key challenge.
    • Existing studies have explored using external structures (e.g., chain structures or narrative diagrams) to enhance the controllability of AI-generated content, but systematic design methods to integrate these structures are lacking.
  • Significance:

    • The creation of complex content such as videos, writing, and music requires dynamic and synchronized tools to ensure smooth human-AI collaboration.
    • Efficient and transparent collaboration tools can significantly improve creative efficiency while reducing the complexity of content generation and management.
  • Research Motivation and Related Work:

    • Related research indicates that using "compositional structures" can effectively organize content elements and facilitate content inspection and iteration.
    • For instance, narrative diagrams can help verify the flow of storylines, while multi-track timelines can support the integration and adjustment of video clips.
    • While there has been research on the individual application of these structures, there is a significant lack of work on effectively integrating multiple structures in complex environments.

Solution

Methodology and Core Concepts

  • Proposed Approach:
    • Introduce compositional structures as the core design element for human-AI co-creation, enabling the visualization and organization of key content components (e.g., spatial, temporal, narrative, and consistency dimensions).
    • Integrate AI within and across these structures to support content generation and automated synchronization, enabling more flexible creative workflows.

Innovation

  • A systematic four-step design method:
    1. Identify the compositional structures and their interrelations required for the creative activity.
    2. Design each individual structure tailored to specific content needs.
    3. Integrate these structures into a unified working environment while defining their synchronization mechanisms.
    4. Embed AI capabilities to optimize the creative process through intelligent generation and association within and across structures.

Implementation Steps and Key Technologies

  1. Defining Structures and Organizational Rules:

    • Identify key dimensions in the content, such as spatial layout, temporal allocation, and narrative flow.
    • Design personalized structures to optimize specific dimensions, such as freeform canvases for material collection and linear text editors for narrative development.
  2. Integration and Synchronization:

    • Use logical mapping and automation techniques to create associations between content across different structures, such as transformations between text and timelines.
    • Provide visual synchronization features, such as highlighting related content or automatically updating linked results.
  3. Embedding AI Capabilities:

    • Within structures, AI assists in generating specific content (e.g., textual narratives or visual scenes).
    • In cross-structure interactions, AI automatically infers and synchronizes content (e.g., adjusting text based on timeline changes to ensure consistency).
  4. Case Application:

    • Developed a video co-creation environment called VideOrigami, incorporating a freeform canvas, linear text editor, grid-based scene planner, and timeline editor.
    • Integrated OpenAI's GPT-4 and DALL-E 3 for text and visual generation to simplify video content creation.

Research Outcomes

Achievements

  1. User Evaluations Confirmed Effectiveness:

    • In user studies, participants reported that the integrated structures helped maintain a clear sense of direction during the creative process and provided greater control through transparent AI generation processes.
    • The scene planner was widely regarded as a critical interaction point, more frequently used as a creative space compared to previous tools.
    • Generation and synchronization features significantly shortened the cycle from a blank canvas to a rough-cut video.
  2. Inspiration for New Workflows:

    • Users exhibited a trend of more frequent cross-structure switching and parallel development workflows.
    • Lower generation costs encouraged users to extensively utilize AI for initial designs, but high evaluation costs hindered large-scale iterations.
  3. Trust and Acceptance of AI-Generated Content:

    • The use of compositional structures allowed participants to intuitively inspect each generation step, fostering stronger trust in AI outcomes.

Advantages and Limitations

  • Comparison with Existing Solutions:

    • Compared to traditional linear workflow systems (e.g., timeline-only editors), compositional structures and synchronization enable more flexible collaboration, reducing the burden of context switching.
    • Unique synchronization capabilities effectively address information loss during content transformations.
  • Limitations:

    • The high fidelity of AI-generated content may restrict users' creative exploration in early stages.
    • Evaluation costs (e.g., reviewing video content) remain high, limiting users' willingness to attempt large-scale changes.

Experimental and Evaluation Results

  • User experience studies showed that AI-assisted synchronization and generation capabilities significantly improved creative efficiency.
  • Both expert and novice users gave consistently high overall ratings for the system, though subtle differences emerged in their experiences of creativity and transparency:
    • Expert users focused more on personalized control, while novice users leaned toward relying on generation features.

Future Directions

  • Explore low-fidelity content generation methods to enhance exploratory potential in early design stages.
  • Develop more efficient tools for evaluating generated content to reduce users' cognitive burden.
  • Investigate how similar methods can be generalized to other complex creative domains (e.g., game design or virtual reality content).

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189208/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713401
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Creative Collaboration & Feedback Systems
work
Professions
Musicians, DJs & Sound Designers, Film & Animation Producers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers