XCreation: A Graph-Based Crossmodal Generative Creativity Support Tool
Authors
Title of the Paper
XCreation: A Graph-Based Crossmodal Generative Creativity Support Tool
Paper Information
- Subject Area: Human-Computer Interaction and Crossmodal Creativity Support Tools
- Keywords: Creativity Support Tools, Crossmodal, Graph Structure, Generative AI, Text-to-Image Generation, Controllable Generation, Graph Interpretability
Research Background and Problem
-
What problems or challenges did the authors identify?
- Most existing creativity support tools (CSTs) only support a single modality (e.g., text or image creation), while crossmodal creativity support remains underexplored.
- CST tools based on generative AI often exhibit "black-box" characteristics, making it difficult for users to understand the generation process, which reduces controllability and interpretability.
- Complex content generation in multimodal creation (e.g., image generation involving multiple entities) faces quality degradation and limitations.
-
Why is this problem important?
- Crossmodal capabilities can more effectively stimulate users' cognitive processes, enhancing creative activities.
- The combination of text and images provides users with unprecedented creative potential and is applicable to various scenarios (e.g., creating storybooks, interactive story illustrations).
-
Research Motivation and Related Work
- Previous studies have shown that multimodal creativity tools can enhance users' creative processes and have theoretically validated the value of crossmodal tools.
- From a neuroscience perspective, crossmodal capabilities promote collaboration between the brain's hemispheres, stimulating more creative possibilities.
- Graphical entity-relation representations can provide better user controllability and logical coherence.
Solution
-
What methods or solutions did the authors propose?
- Proposed a crossmodal creativity support tool, XCreation, which combines generative AI to support multimodal transformations such as text-to-image and graph-to-text.
- Utilized interpretable entity-relation graphs to visually represent elements and relationships in images, improving the controllability and transparency of the generation process.
- Provided an advanced multimodal creation interface, including text, graph, and hand-drawing functionalities.
-
What are the innovative aspects of this solution?
- Introduced graph structures to help users decompose complex ideas into atomic elements, gradually expanding their creativity.
- Used graphs as intermediate representations, enabling seamless modality switching and efficient idea transfer.
- Offered fine-grained element control (e.g., modifying a specific object in an image without affecting other elements).
-
What are the implementation steps and key technologies used?
- Text Input: Users input a storyline, and the system uses natural language processing tools (e.g., SpaCy and GPT-3) to extract entities and relationships, generating corresponding graph structures and initial images.
- Graph Editing: Users can drag, adjust, add, or delete nodes and edges in the graph, while generating corresponding image elements.
- Graph-to-Text: Generate natural language sentences based on the graph (text enhancement supported by GPT-3).
- Hand-Drawn Input: Users can create custom drawings, which are converted into graph nodes.
- Final Generation: Based on the user-designed graph canvas, complete images are generated using generative AI models (e.g., DALL-E 2).
Research Outcomes
-
What specific outcomes were achieved?
- User studies demonstrated that XCreation effectively enhanced creativity, controllability, and interpretability in story creation.
- Compared to single-modality baseline tools, XCreation significantly improved creation efficiency and supported higher levels of controllability.
- The introduction of graph structures notably enhanced the logical coherence of the creative process and encouraged users to decompose ideas.
-
What advantages does it have over existing solutions?
- Supports more natural crossmodal switching processes.
- The graphical interface facilitates creative expression and improves the quality of complex content generation.
- Combines generative AI with graphical methods, making the creation process more intuitive and transparent.
-
What were the experimental or evaluation results?
- User satisfaction with the XCreation system (SUS score): 85.5, indicating high usability and user approval.
- Quantitative results showed that XCreation outperformed baseline tools (AI-Storyteller and DALL-E 2) in terms of creative expression, adjustability, and multimodal switching.
- Experimental tasks, including closed-ended creation, re-creation, and open-ended creation, consistently demonstrated its advantages.
-
Limitations and Future Directions
-
Limitations:
- Node layout relies entirely on manual adjustments, which may increase users' cognitive load.
- The association between graphs and image generation is limited to the node level, without direct edge-based content generation.
- Usability for long logical stories and visual resources still needs improvement.
-
Future Directions:
- Introduce automatic layout functionality to reduce users' time investment using graph theory algorithms.
- Expand support for video animation generation, providing a more vivid digital storytelling mode.
- Integrate additional modalities (e.g., music or sound) to further enhance multimodal creation possibilities.
- Develop crossmodal content commercial templates to enhance practical application value.
-
XCreation establishes an efficient and transparent paradigm for creativity support tools, offering valuable design insights for innovations in human-computer interaction and generative AI in the creative domain.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can cross-modal creative support tools enhance users' creativity and cognitive processes?Category: GenAI Creative Control and Co-CreationSimilar questionsarrow_forward
- Can graph structures provide greater creative control and logical coherence?Category: GenAI Creative Control and Co-CreationSimilar questionsarrow_forward
- How can high-quality content creation among multiple entities be achieved in image generation?Category: GenAI Creative Control and Co-CreationSimilar questionsarrow_forward
Practical Problems
1- Users struggle to intuitively control the generation process when creating across text and images.Category: GenAI Creative Control and Co-CreationSimilar questionsarrow_forward
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)