Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Image-Text Co-Editing
Authors
Paper Title
Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Image-Text Co-Editing
Publication Info
- Topic area: Multimodal systems for creative writing, focusing on image-text co-editing.
- Keywords: Multimodal interaction, fictional story writing, image-text alignment, creativity support, instrumental interaction, co-editing, narrative development, LLMs, cognitive processes, user studies.
Background and Problem
- Problem / challenge: Existing writing tools are predominantly text-centric, treating visuals as supplementary rather than integral to the creative process. This separation increases cognitive load and limits the integration of visual and textual elements in narrative development.
- Significance: Fictional story writing is inherently multimodal, relying on both visual and textual channels for ideation and expression. A system that integrates these modalities could enhance creativity, reduce cognitive load, and improve narrative coherence.
- Motivation and related work: Previous tools for creative writing have used visuals for inspiration or structural organization but lack mechanisms for synchronized, manipulable image-text editing. Recent multimodal systems and LLM-powered tools have shown promise but still treat images and text as separate entities. This paper addresses the gap by designing a system that tightly integrates these modalities.
Solution
- Proposed approach: Vistoria, a multimodal system that enables synchronized image-text co-editing through instrumental operations, facilitating narrative exploration and development.
- Novelty:
- Introduction of instrumental operations (Lasso, Collage, Perspective Shift, Filter) for synchronized image-text editing.
- A cyclic workflow that aligns visual and textual elements, enabling iterative narrative development.
- A Cluster panel for organizing and reusing narrative fragments into coherent storylines.
- Procedure and key techniques:
- Multimodal inputs (text, sketches, images) are transformed into image-text cards.
- Instrumental operations allow users to manipulate and align text and images simultaneously.
- A Cluster panel aggregates narrative elements for structured organization.
- The system employs a multi-agent backend for narrative construction, visual synthesis, and memory management.
Results
- Concrete findings:
- Vistoria enhanced expressiveness (mean score: 6.08 vs. 4.33 in baseline, p=0.023) and immersion (mean score: 4.92 vs. 2.75, p=0.0006).
- Participants explored more directions (mean: 6.92 vs. 1.42) and branches (mean: 3.00 vs. 1.92) compared to the baseline.
- Increased mental (mean: 5.16 vs. 3.17, p=0.0000) and physical workload (mean: 4.67 vs. 2.08, p=0.0002) was observed.
- Advantage over baselines:
- Supported divergent exploration and richer narrative development.
- Preserved participants' sense of agency and ownership over the creative process.
- Enabled synchronized image-text editing, unlike text-only GPT workflows.
- Experiments / evaluation:
- Conducted a controlled user study with 12 participants (aged 21–32, all with creative writing experience).
- Compared Vistoria against a text-only GPT baseline using NASA-TLX and Creativity Support Index (CSI) surveys, along with qualitative interviews and system logs.
- Limitations and future work:
- Short-term tasks (300–500 words) may not reflect long-form writing demands.
- Small, culturally homogeneous participant sample limits generalizability.
- Higher cognitive load for first-time users; future designs could include adaptive features and support for other modalities like audio.
Summary
Vistoria is a multimodal system designed to support fictional story writing by integrating synchronized image-text co-editing. It introduces instrumental operations (Lasso, Collage, Perspective Shift, Filter) to enable fluid transitions between modalities, enhancing expressiveness, immersion, and exploratory ideation. While the system increases cognitive demand, it preserves users' sense of agency and ownership, treating them as active creators rather than passive editors. Future work will address scalability for long-form writing, diverse participant inclusion, and adaptive features to reduce workload.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
Sci-Fi Spark: A Human-AI Co-Creation System for Science Fiction Ideation
CHI '26· AI-Assisted Creative Writing +2
- 67%
Where Do I 'Add the Egg'?: Exploring Agency and Ownership in AI Creative Co-Writing Systems
CHI '26· AI-Assisted Creative Writing +2
Based on Jaccard similarity of research subtopics & professions (≥60%)