XCreation: A Graph-Based Crossmodal Generative Creativity Support Tool

Generative AI (Text, Image, Music, Video)Human-LLM CollaborationExplainable AI (XAI)Interactive Narrative & Immersive StorytellingContent Creators (YouTubers, Podcasters)Film & Animation ProducersUI/UX DesignersVisual Artists & Designers

Title of the Paper

XCreation: A Graph-Based Crossmodal Generative Creativity Support Tool

Paper Information

  • Subject Area: Human-Computer Interaction and Crossmodal Creativity Support Tools
  • Keywords: Creativity Support Tools, Crossmodal, Graph Structure, Generative AI, Text-to-Image Generation, Controllable Generation, Graph Interpretability

Research Background and Problem

  • What problems or challenges did the authors identify?

    • Most existing creativity support tools (CSTs) only support a single modality (e.g., text or image creation), while crossmodal creativity support remains underexplored.
    • CST tools based on generative AI often exhibit "black-box" characteristics, making it difficult for users to understand the generation process, which reduces controllability and interpretability.
    • Complex content generation in multimodal creation (e.g., image generation involving multiple entities) faces quality degradation and limitations.
  • Why is this problem important?

    • Crossmodal capabilities can more effectively stimulate users' cognitive processes, enhancing creative activities.
    • The combination of text and images provides users with unprecedented creative potential and is applicable to various scenarios (e.g., creating storybooks, interactive story illustrations).
  • Research Motivation and Related Work

    • Previous studies have shown that multimodal creativity tools can enhance users' creative processes and have theoretically validated the value of crossmodal tools.
    • From a neuroscience perspective, crossmodal capabilities promote collaboration between the brain's hemispheres, stimulating more creative possibilities.
    • Graphical entity-relation representations can provide better user controllability and logical coherence.

Solution

  • What methods or solutions did the authors propose?

    • Proposed a crossmodal creativity support tool, XCreation, which combines generative AI to support multimodal transformations such as text-to-image and graph-to-text.
    • Utilized interpretable entity-relation graphs to visually represent elements and relationships in images, improving the controllability and transparency of the generation process.
    • Provided an advanced multimodal creation interface, including text, graph, and hand-drawing functionalities.
  • What are the innovative aspects of this solution?

    • Introduced graph structures to help users decompose complex ideas into atomic elements, gradually expanding their creativity.
    • Used graphs as intermediate representations, enabling seamless modality switching and efficient idea transfer.
    • Offered fine-grained element control (e.g., modifying a specific object in an image without affecting other elements).
  • What are the implementation steps and key technologies used?

    1. Text Input: Users input a storyline, and the system uses natural language processing tools (e.g., SpaCy and GPT-3) to extract entities and relationships, generating corresponding graph structures and initial images.
    2. Graph Editing: Users can drag, adjust, add, or delete nodes and edges in the graph, while generating corresponding image elements.
    3. Graph-to-Text: Generate natural language sentences based on the graph (text enhancement supported by GPT-3).
    4. Hand-Drawn Input: Users can create custom drawings, which are converted into graph nodes.
    5. Final Generation: Based on the user-designed graph canvas, complete images are generated using generative AI models (e.g., DALL-E 2).

Research Outcomes

  • What specific outcomes were achieved?

    • User studies demonstrated that XCreation effectively enhanced creativity, controllability, and interpretability in story creation.
    • Compared to single-modality baseline tools, XCreation significantly improved creation efficiency and supported higher levels of controllability.
    • The introduction of graph structures notably enhanced the logical coherence of the creative process and encouraged users to decompose ideas.
  • What advantages does it have over existing solutions?

    • Supports more natural crossmodal switching processes.
    • The graphical interface facilitates creative expression and improves the quality of complex content generation.
    • Combines generative AI with graphical methods, making the creation process more intuitive and transparent.
  • What were the experimental or evaluation results?

    • User satisfaction with the XCreation system (SUS score): 85.5, indicating high usability and user approval.
    • Quantitative results showed that XCreation outperformed baseline tools (AI-Storyteller and DALL-E 2) in terms of creative expression, adjustability, and multimodal switching.
    • Experimental tasks, including closed-ended creation, re-creation, and open-ended creation, consistently demonstrated its advantages.
  • Limitations and Future Directions

    • Limitations:

      • Node layout relies entirely on manual adjustments, which may increase users' cognitive load.
      • The association between graphs and image generation is limited to the node level, without direct edge-based content generation.
      • Usability for long logical stories and visual resources still needs improvement.
    • Future Directions:

      • Introduce automatic layout functionality to reduce users' time investment using graph theory algorithms.
      • Expand support for video animation generation, providing a more vivid digital storytelling mode.
      • Integrate additional modalities (e.g., music or sound) to further enhance multimodal creation possibilities.
      • Develop crossmodal content commercial templates to enhance practical application value.

XCreation establishes an efficient and transparent paradigm for creativity support tools, offering valuable design insights for innovations in human-computer interaction and generative AI in the creative domain.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/126857/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3586183.3606826
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration, Explainable AI (XAI), Interactive Narrative & Immersive Storytelling
work
Professions
Content Creators (YouTubers, Podcasters), Film & Animation Producers, UI/UX Designers, Visual Artists & Designers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers