WorldSmith: A Multi-Modal Image Synthesis Tool for Fictional World Building

Generative AI (Text, Image, Music, Video)AI-Assisted Creative WritingGraphic Design & Typography ToolsVisual Artists & DesignersFreelancers (Design, Writing, Translation)

Document Title

WorldSmith: Iterative and Expressive Prompting for World Building with a Generative AI

Document Information

  • Domain: Human-Computer Interaction, AI Generation, and Virtual World Building
  • Keywords: Multimodal Image Generation, Virtual World Building, AI-Assisted Creativity, Layered Editing, Expressive Prompting

Research Background and Problem

  • Challenge: Constructing virtual worlds requires complex and detailed visual representation, which can be difficult for creators lacking artistic skills or experience. Existing generative image models like Dall-E and Stable Diffusion typically use "one-click" description interfaces, which fail to meet the iterative needs of virtual world building.
  • Importance: In entertainment formats such as games and films, a meticulously crafted virtual world is crucial for success. Rapid and effective visualization of worlds can accelerate collaboration among creators and inspire creativity.
  • Research Motivation: To explore expressive prompting tools based on generative AI that enable users to progressively create complex visual worlds using multimodal inputs such as text, sketches, and region-based drawings.

Solution

  • Approach: WorldSmith is a multimodal generative tool specifically designed for virtual world building. It starts from a base and iteratively generates and refines virtual worlds through layered image editing and interactive visual prompts.
  • Innovations:
    • Supports text, sketches, and region masks as inputs, allowing spatial prompts through non-text methods.
    • Introduces layered generation and editing, enabling users to refine virtual worlds progressively from global image composition to local details.
    • Provides a tree-view to trace user interaction history, supporting exploration of different iterative paths.
  • Implementation Steps:
    1. Text Input: Users provide global scene descriptions or detailed specifications for specific areas.
    2. Sketching and Masking: Users draw rough sketches or designate regions to provide spatial prompts to the model.
    3. Image Generation and Editing: Images are generated based on user-defined inputs, with iterative improvement options, including multi-image composition.
    4. Tree-View: Records user generation history, enabling backtracking or branching into alternative versions.

Research Outcomes

  • Specific Results:
    • WorldSmith enables users to interact with prompt-based generative models in a more engaging manner, surpassing the limitations of "one-click" interactions.
    • A user study involving 13 participants validated the tool's practicality and efficiency in virtual world building, helping creators quickly produce initial drafts.
    • A total of 2,748 images were generated, with 86 image composition operations completed.
  • Advantages:
    • Supports multimodal inputs like sketching and region segmentation, enhancing users' ability to express their intentions.
    • Emphasizes layered generation and editing, facilitating a workflow from rough drafts to detailed creations.
    • Simulates the iterative nature of virtual world building, increasing users' sense of control over generative AI.
  • Experiment and Evaluation Results:
    • Users rated the tool's usability highly overall. Surveys and behavioral records revealed a tendency to first generate drafts and then refine details iteratively.
    • The tree-view interaction model positively impacted the tracing of generation history and exploration of branches.
    • Almost all participants preferred multimodal input formats over single-text prompts.
  • Limitations and Future Directions:
    • Image composition results may face potential issues with style and perspective inconsistency. More precise post-processing mechanisms like MultiDiffusion could be introduced.
    • For more complex scenes and asset generation, users might require additional time or technical support.
    • Other dimensions of virtual world building, such as narrative context, music design, or character creation, remain unaddressed.

Conclusion

WorldSmith surpasses the "one-click" interaction limitations of existing generative image tools by introducing layered generation and spatial prompts, providing creators with a more expressive and iterative tool. The research outcomes not only demonstrate the tool's effectiveness in virtual world building but also propose "layered prompting" and "spatial prompting" as new directions for broader AI prompt interaction design.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/126668/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3586183.3606772
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), AI-Assisted Creative Writing, Graphic Design & Typography Tools
work
Professions
Visual Artists & Designers, Freelancers (Design, Writing, Translation)
article
Content Status
Full text indexed
hub
Related Papers
4 related papers