Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Image-Text Co-Editing

AI-Assisted Creative WritingCreative Collaboration & Feedback SystemsContent Creators (YouTubers, Podcasters)HCI Researchers

Paper Title

Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Image-Text Co-Editing

Publication Info

  • Topic area: Multimodal systems for creative writing, focusing on image-text co-editing.
  • Keywords: Multimodal interaction, fictional story writing, image-text alignment, creativity support, instrumental interaction, co-editing, narrative development, LLMs, cognitive processes, user studies.

Background and Problem

  • Problem / challenge: Existing writing tools are predominantly text-centric, treating visuals as supplementary rather than integral to the creative process. This separation increases cognitive load and limits the integration of visual and textual elements in narrative development.
  • Significance: Fictional story writing is inherently multimodal, relying on both visual and textual channels for ideation and expression. A system that integrates these modalities could enhance creativity, reduce cognitive load, and improve narrative coherence.
  • Motivation and related work: Previous tools for creative writing have used visuals for inspiration or structural organization but lack mechanisms for synchronized, manipulable image-text editing. Recent multimodal systems and LLM-powered tools have shown promise but still treat images and text as separate entities. This paper addresses the gap by designing a system that tightly integrates these modalities.

Solution

  • Proposed approach: Vistoria, a multimodal system that enables synchronized image-text co-editing through instrumental operations, facilitating narrative exploration and development.
  • Novelty:
    1. Introduction of instrumental operations (Lasso, Collage, Perspective Shift, Filter) for synchronized image-text editing.
    2. A cyclic workflow that aligns visual and textual elements, enabling iterative narrative development.
    3. A Cluster panel for organizing and reusing narrative fragments into coherent storylines.
  • Procedure and key techniques:
    • Multimodal inputs (text, sketches, images) are transformed into image-text cards.
    • Instrumental operations allow users to manipulate and align text and images simultaneously.
    • A Cluster panel aggregates narrative elements for structured organization.
    • The system employs a multi-agent backend for narrative construction, visual synthesis, and memory management.

Results

  • Concrete findings:
    • Vistoria enhanced expressiveness (mean score: 6.08 vs. 4.33 in baseline, p=0.023) and immersion (mean score: 4.92 vs. 2.75, p=0.0006).
    • Participants explored more directions (mean: 6.92 vs. 1.42) and branches (mean: 3.00 vs. 1.92) compared to the baseline.
    • Increased mental (mean: 5.16 vs. 3.17, p=0.0000) and physical workload (mean: 4.67 vs. 2.08, p=0.0002) was observed.
  • Advantage over baselines:
    • Supported divergent exploration and richer narrative development.
    • Preserved participants' sense of agency and ownership over the creative process.
    • Enabled synchronized image-text editing, unlike text-only GPT workflows.
  • Experiments / evaluation:
    • Conducted a controlled user study with 12 participants (aged 21–32, all with creative writing experience).
    • Compared Vistoria against a text-only GPT baseline using NASA-TLX and Creativity Support Index (CSI) surveys, along with qualitative interviews and system logs.
  • Limitations and future work:
    • Short-term tasks (300–500 words) may not reflect long-form writing demands.
    • Small, culturally homogeneous participant sample limits generalizability.
    • Higher cognitive load for first-time users; future designs could include adaptive features and support for other modalities like audio.

Summary

Vistoria is a multimodal system designed to support fictional story writing by integrating synchronized image-text co-editing. It introduces instrumental operations (Lasso, Collage, Perspective Shift, Filter) to enable fluid transitions between modalities, enhancing expressiveness, immersion, and exploratory ideation. While the system increases cognitive demand, it preserves users' sense of agency and ownership, treating them as active creators rather than passive editors. Future work will address scalability for long-form writing, diverse participant inclusion, and adaptive features to reduce workload.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222530/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790400
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
AI-Assisted Creative Writing, Creative Collaboration & Feedback Systems
work
Professions
Content Creators (YouTubers, Podcasters), HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
2 related papers