Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
Generative AI (Text, Image, Music, Video)AI-Assisted Creative WritingInteractive Narrative & Immersive StorytellingGame Developers & DesignersFilm & Animation ProducersVisual Artists & Designers
Research Background and Issues
- Issues and Challenges: The authors observed that many AI storytelling tools primarily rely on natural language input to generate story content, neglecting other possible interaction modes, such as storytelling activities conducted by children through toy play. Additionally, existing AI technologies remain limited in supporting multimodal inputs, such as lacking intuitive understanding of simplified inputs, high interaction latency, or requiring users to manipulate complex input forms.
- Significance: Storytelling is an important form of creative expression. By developing multimodal interaction technologies, it is possible to expand the potential for human-AI collaboration in creative expression, improve user experience, and create new research opportunities in technology and interaction design.
- Motivation and Related Work: This study is inspired by Heider and Simmel's experiments on the actions of shapes, which can convey rich social interactions. Moreover, past research has attempted to use AI to generate story text, worlds, or visual content, but these efforts have mostly focused on single modalities (e.g., text) rather than multimodal interactions.
Solution
- Method or Solution: The authors propose a multimodal AI interaction system called Toyteller, which combines toy manipulation and story generation. It uses action symbols as input to guide story generation, while also incorporating these actions as part of the story output.
- Innovations:
- Introducing "toy manipulation" as a novel input mode, allowing users to express storytelling intentions through symbolic actions.
- Leveraging a conversion layer between actions and text to enable effective operation of AI in both action generation and text generation.
- Implementation Steps:
- Develop a toy manipulation interface that allows users to physically control symbolic actions or provide text input.
- Train an AI model to recognize events and active character information from actions and generate text and actions based on this information.
- Design a user-friendly tool interface to support story setup, action recording, text generation, and interaction flexibility.
- Key Technologies:
- Using LSTM for action sequence processing to optimize speed.
- Embedding action information into text embedding space to achieve multimodal conversion.
- Applying soft prompting techniques to optimize the generation performance of large language models.
Research Outcomes
- Specific Outcomes:
- Toyteller outperformed GPT-4o in tasks such as action-to-story text generation, action generation, and action recognition, particularly in terms of accuracy, novelty of generation, and interaction smoothness.
- User studies revealed that toy interaction is well-suited for expressing ambiguous intentions that are difficult to describe through language and is significantly helpful in inspiring creativity.
- Advantages Comparison:
- Compared to GPT-4o, Toyteller has lower interaction latency, higher action recognition accuracy, and greater novelty in story text generation.
- The system allows greater interaction flexibility, enabling users to freely allocate roles between AI and themselves in the creative process.
- Experimental Results:
- In action recognition tests, Toyteller significantly improved correct action ranking and weight allocation while markedly reducing processing latency.
- In text generation evaluations, Toyteller's content was more creative and highly consistent with the actions.
- In action generation tasks, Toyteller produced actions that were more realistic and aligned with target event requirements.
- Limitations and Future Directions:
- Limitations: In certain scenarios, the system may produce biases when interpreting actions and has limited understanding of specific domains (e.g., sports actions). The current "triangle" symbol design restricts visual expression.
- Future Directions:
- Expand scenario complexity (e.g., adding more characters and props).
- Improve visual presentation (e.g., adding colors and more complex symbols).
- Integrate other interaction forms, such as 3D actions, sound, or tactile input.
- Enhance adaptability to professional domains and complex scenarios.
Conclusion
Toyteller elevates AI-powered storytelling to the realm of multimodal interaction through its innovative toy interaction approach. This study not only demonstrates technological innovation but also highlights the potential of this interaction mode in creative expression, tool design, and user experience. Future work could further optimize technical performance, expand application scenarios, and explore possibilities for integration with other generative technologies.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- Can introducing toy manipulation as an input mode enhance multimodal interaction in AI storytelling tools?Category: Generative Writing and Storytelling ControlSimilar questionsarrow_forward
- How can user behavioral symbols be efficiently mapped into text embedding space during story generation?Category: Generative Writing and Storytelling ControlSimilar questionsarrow_forward
- Can toy interaction effectively express vague intents hard to describe in language and inspire user creativity?Category: Generative Writing and Storytelling ControlSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Existing AI storytelling tools poorly support multimodal input, with high interaction latency and complex forms.Category: Input Performance, Accidental Touch Control, and Interaction EfficiencySimilar questionsarrow_forward
- 83%
WhatELSE: Shaping Narrative Spaces at Configurable Level of Abstraction for AI-bridged Interactive Storytelling
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 67%
Generative AI in Documentary Photography: Exploring Opportunities and Challenges for Visual Storytelling
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 67%
WhatIF: Branched Narrative Fiction Visualization for Authoring Emergent Narratives using Large Language Models
C&C '25· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713435
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), AI-Assisted Creative Writing, Interactive Narrative & Immersive Storytelling
work
Professions
Game Developers & Designers, Film & Animation Producers, Visual Artists & Designers
article
Content Status
Full text indexed
hub
Related Papers
3 related papers