Compositional Structures as Substrates for Human-AI Co-creation Environment: A Design Approach and A Case Study
Authors
Generative AI (Text, Image, Music, Video)Creative Collaboration & Feedback SystemsMusicians, DJs & Sound DesignersFilm & Animation ProducersUI/UX Designers
Research Background and Issues
-
Issues and Challenges:
- Current human-AI co-creation environments often rely on a linear "prompt-result" interaction model, lacking support for exploration, planning, iteration, and control or inspection of AI-generated content during the creative process.
- In complex content creation processes, such as video production, effectively organizing and managing multidimensional information (e.g., narrative, temporal, spatial classification) remains a key challenge.
- Existing studies have explored using external structures (e.g., chain structures or narrative diagrams) to enhance the controllability of AI-generated content, but systematic design methods to integrate these structures are lacking.
-
Significance:
- The creation of complex content such as videos, writing, and music requires dynamic and synchronized tools to ensure smooth human-AI collaboration.
- Efficient and transparent collaboration tools can significantly improve creative efficiency while reducing the complexity of content generation and management.
-
Research Motivation and Related Work:
- Related research indicates that using "compositional structures" can effectively organize content elements and facilitate content inspection and iteration.
- For instance, narrative diagrams can help verify the flow of storylines, while multi-track timelines can support the integration and adjustment of video clips.
- While there has been research on the individual application of these structures, there is a significant lack of work on effectively integrating multiple structures in complex environments.
Solution
Methodology and Core Concepts
- Proposed Approach:
- Introduce compositional structures as the core design element for human-AI co-creation, enabling the visualization and organization of key content components (e.g., spatial, temporal, narrative, and consistency dimensions).
- Integrate AI within and across these structures to support content generation and automated synchronization, enabling more flexible creative workflows.
Innovation
- A systematic four-step design method:
- Identify the compositional structures and their interrelations required for the creative activity.
- Design each individual structure tailored to specific content needs.
- Integrate these structures into a unified working environment while defining their synchronization mechanisms.
- Embed AI capabilities to optimize the creative process through intelligent generation and association within and across structures.
Implementation Steps and Key Technologies
-
Defining Structures and Organizational Rules:
- Identify key dimensions in the content, such as spatial layout, temporal allocation, and narrative flow.
- Design personalized structures to optimize specific dimensions, such as freeform canvases for material collection and linear text editors for narrative development.
-
Integration and Synchronization:
- Use logical mapping and automation techniques to create associations between content across different structures, such as transformations between text and timelines.
- Provide visual synchronization features, such as highlighting related content or automatically updating linked results.
-
Embedding AI Capabilities:
- Within structures, AI assists in generating specific content (e.g., textual narratives or visual scenes).
- In cross-structure interactions, AI automatically infers and synchronizes content (e.g., adjusting text based on timeline changes to ensure consistency).
-
Case Application:
- Developed a video co-creation environment called VideOrigami, incorporating a freeform canvas, linear text editor, grid-based scene planner, and timeline editor.
- Integrated OpenAI's GPT-4 and DALL-E 3 for text and visual generation to simplify video content creation.
Research Outcomes
Achievements
-
User Evaluations Confirmed Effectiveness:
- In user studies, participants reported that the integrated structures helped maintain a clear sense of direction during the creative process and provided greater control through transparent AI generation processes.
- The scene planner was widely regarded as a critical interaction point, more frequently used as a creative space compared to previous tools.
- Generation and synchronization features significantly shortened the cycle from a blank canvas to a rough-cut video.
-
Inspiration for New Workflows:
- Users exhibited a trend of more frequent cross-structure switching and parallel development workflows.
- Lower generation costs encouraged users to extensively utilize AI for initial designs, but high evaluation costs hindered large-scale iterations.
-
Trust and Acceptance of AI-Generated Content:
- The use of compositional structures allowed participants to intuitively inspect each generation step, fostering stronger trust in AI outcomes.
Advantages and Limitations
-
Comparison with Existing Solutions:
- Compared to traditional linear workflow systems (e.g., timeline-only editors), compositional structures and synchronization enable more flexible collaboration, reducing the burden of context switching.
- Unique synchronization capabilities effectively address information loss during content transformations.
-
Limitations:
- The high fidelity of AI-generated content may restrict users' creative exploration in early stages.
- Evaluation costs (e.g., reviewing video content) remain high, limiting users' willingness to attempt large-scale changes.
Experimental and Evaluation Results
- User experience studies showed that AI-assisted synchronization and generation capabilities significantly improved creative efficiency.
- Both expert and novice users gave consistently high overall ratings for the system, though subtle differences emerged in their experiences of creativity and transparency:
- Expert users focused more on personalized control, while novice users leaned toward relying on generation features.
Future Directions
- Explore low-fidelity content generation methods to enhance exploratory potential in early design stages.
- Develop more efficient tools for evaluating generated content to reduce users' cognitive burden.
- Investigate how similar methods can be generalized to other complex creative domains (e.g., game design or virtual reality content).
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can compositional structures (e.g., narrative graphs, sequence diagrams) improve efficiency and transparency in human-AI collaboration for complex content creation?Category: Data Storytelling and Narrative Visualization NeedsSimilar questionsarrow_forward
- In complex content decomposition, how can multiple structures be designed and integrated to support exploration, iteration, and control?Category: Data Storytelling and Narrative Visualization NeedsSimilar questionsarrow_forward
- How can AI efficiently generate, relate, and synchronize content across structural interactions?Category: Data Storytelling and Narrative Visualization NeedsSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Creators struggle to effectively organize and manage content across multiple dimensions in complex scenarios such as video production.Category: Data Storytelling and Narrative Visualization NeedsSimilar questionsarrow_forward
- 83%
SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for Video
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 67%
Sound Designer-Generative AI Interactions: Towards Designing Creative Support Tools for Professional Sound Designers
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 67%
Understanding User Perceptions, Collaborative Experience and User Engagement in Different Human-AI Interaction Designs for Co-Creative Systems
C&C '22· Generative AI (Text, Image, Music, Video) +2
- 67%
Exploring the Potential for Generative AI-based Conversational Cues for Real-Time Collaborative Ideation
C&C '24· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713401
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Creative Collaboration & Feedback Systems
work
Professions
Musicians, DJs & Sound Designers, Film & Animation Producers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers