WorldSmith: A Multi-Modal Image Synthesis Tool for Fictional World Building
Authors
Document Title
WorldSmith: Iterative and Expressive Prompting for World Building with a Generative AI
Document Information
- Domain: Human-Computer Interaction, AI Generation, and Virtual World Building
- Keywords: Multimodal Image Generation, Virtual World Building, AI-Assisted Creativity, Layered Editing, Expressive Prompting
Research Background and Problem
- Challenge: Constructing virtual worlds requires complex and detailed visual representation, which can be difficult for creators lacking artistic skills or experience. Existing generative image models like Dall-E and Stable Diffusion typically use "one-click" description interfaces, which fail to meet the iterative needs of virtual world building.
- Importance: In entertainment formats such as games and films, a meticulously crafted virtual world is crucial for success. Rapid and effective visualization of worlds can accelerate collaboration among creators and inspire creativity.
- Research Motivation: To explore expressive prompting tools based on generative AI that enable users to progressively create complex visual worlds using multimodal inputs such as text, sketches, and region-based drawings.
Solution
- Approach: WorldSmith is a multimodal generative tool specifically designed for virtual world building. It starts from a base and iteratively generates and refines virtual worlds through layered image editing and interactive visual prompts.
- Innovations:
- Supports text, sketches, and region masks as inputs, allowing spatial prompts through non-text methods.
- Introduces layered generation and editing, enabling users to refine virtual worlds progressively from global image composition to local details.
- Provides a tree-view to trace user interaction history, supporting exploration of different iterative paths.
- Implementation Steps:
- Text Input: Users provide global scene descriptions or detailed specifications for specific areas.
- Sketching and Masking: Users draw rough sketches or designate regions to provide spatial prompts to the model.
- Image Generation and Editing: Images are generated based on user-defined inputs, with iterative improvement options, including multi-image composition.
- Tree-View: Records user generation history, enabling backtracking or branching into alternative versions.
Research Outcomes
- Specific Results:
- WorldSmith enables users to interact with prompt-based generative models in a more engaging manner, surpassing the limitations of "one-click" interactions.
- A user study involving 13 participants validated the tool's practicality and efficiency in virtual world building, helping creators quickly produce initial drafts.
- A total of 2,748 images were generated, with 86 image composition operations completed.
- Advantages:
- Supports multimodal inputs like sketching and region segmentation, enhancing users' ability to express their intentions.
- Emphasizes layered generation and editing, facilitating a workflow from rough drafts to detailed creations.
- Simulates the iterative nature of virtual world building, increasing users' sense of control over generative AI.
- Experiment and Evaluation Results:
- Users rated the tool's usability highly overall. Surveys and behavioral records revealed a tendency to first generate drafts and then refine details iteratively.
- The tree-view interaction model positively impacted the tracing of generation history and exploration of branches.
- Almost all participants preferred multimodal input formats over single-text prompts.
- Limitations and Future Directions:
- Image composition results may face potential issues with style and perspective inconsistency. More precise post-processing mechanisms like MultiDiffusion could be introduced.
- For more complex scenes and asset generation, users might require additional time or technical support.
- Other dimensions of virtual world building, such as narrative context, music design, or character creation, remain unaddressed.
Conclusion
WorldSmith surpasses the "one-click" interaction limitations of existing generative image tools by introducing layered generation and spatial prompts, providing creators with a more expressive and iterative tool. The research outcomes not only demonstrate the tool's effectiveness in virtual world building but also propose "layered prompting" and "spatial prompting" as new directions for broader AI prompt interaction design.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can multimodal inputs (e.g., text, sketches, region drawing) help users build complex virtual worlds?Category: Creative Workflows and Multi-Stage PipelinesSimilar questionsarrow_forward
- How does the WorldSmith tool support iterative image editing workflows to enhance expressiveness in virtual world building?Category: Creative Workflows and Multi-Stage PipelinesSimilar questionsarrow_forward
- How does a tree-view interaction pattern help users track generation history and explore different creative paths?Category: Creative Workflows and Multi-Stage PipelinesSimilar questionsarrow_forward
Practical Problems
1- Many creators struggle to build detailed virtual worlds due to a lack of artistic skills.Category: Creative Workflows and Multi-Stage PipelinesSimilar questionsarrow_forward
- 80%
PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like Interactions
UIST '23· Generative AI (Text, Image, Music, Video) +1
- 67%
DeepWriting: Making Digital Ink Editable via Deep Generative Modeling
CHI '18· Generative AI (Text, Image, Music, Video) +1
- 60%
The Value, Benefits, and Concerns of Generative AI-Powered Assistance in Writing
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 60%
Lyric Poetry in the Face of Posthumanism: An Analysis of Generative AI-Assisted Poetry Writing
C&C '25· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)