Vidmento: Scaffolded Expansion for Video Storytelling with Generative Video
Authors
Paper Title
Vidmento: Creating Video Stories through Context-Aware Expansion with Generative Video
Publication Info
- Topic area: Hybrid video storytelling using generative and captured media
- Keywords: Generative video, hybrid storytelling, video editing, AI-assisted creation, narrative expansion, cinematic principles, creative tools, video authoring, context-aware generation, generative AI
Background and Problem
- Problem / challenge: Video storytelling is constrained by the availability of captured footage, limiting creators' ability to fill narrative gaps or explore alternative storylines. Existing generative video tools often fail to integrate seamlessly with captured media, lacking contextual blending and narrative coherence.
- Significance: Addressing these limitations enables creators to craft richer, more cohesive video stories by combining the authenticity of captured footage with the flexibility of generative media.
- Motivation and related work: Prior tools focus on either fully captured or fully generated workflows, often treating generative video as an isolated process. This paper identifies a gap in hybrid storytelling approaches and aims to bridge it by blending captured and generative media within a unified framework.
Solution
- Proposed approach: Vidmento, a video authoring tool that implements the "generative expansion" framework to blend captured and generative media for hybrid video storytelling.
- Novelty:
- Introduction of "generative expansion," a design framework for contextually blending captured and generative video.
- Development of Vidmento, a tool that integrates narrative and cinematic principles to support hybrid video creation.
- Empirical findings from formative interviews and a user study highlighting opportunities and challenges in hybrid storytelling.
- Procedure and key techniques:
- Vidmento provides a semi-structured canvas for organizing and sequencing media, coupled with a script editor for narration.
- Context-aware AI generates new video clips and script suggestions to fill narrative gaps or expand stories.
- Users retain creative control through features like annotation-based refinements, structured prompts, and timeline alignment tools.
- The system employs a two-stage pipeline for generating contextual keyframes and animating them into videos.
Results
- Concrete findings:
- Vidmento enabled users to augment their original media by an average of 96.4%, generating ∼31 images and 8 videos per session.
- The system supported narrative exploration, with creators clicking an average of 3.7 story suggestions and generating 2.0 refined script versions per session.
- Advantage over baselines:
- Unlike existing tools, Vidmento integrates captured and generative media within a unified workspace, offering context-aware blending and narrative scaffolding.
- Users valued its ability to visualize and manipulate stories through a semi-structured canvas and synchronized script editor.
- Experiments / evaluation:
- Conducted a user study with 12 creators (novices to professionals) who produced 1–2 minute narrated videos using Vidmento.
- Participants highlighted the system's utility for ideation, narrative expansion, and blending media, though some noted challenges with refining generative outputs.
- Limitations and future work:
- Challenges include style mismatches between captured and generated footage, difficulty refining outputs, and concerns about authenticity in personal stories.
- Future directions include improving model adaptability, exploring spatial and stylistic blending, and designing fully unified multimodal canvases.
Summary
Vidmento introduces a novel framework, "generative expansion," for hybrid video storytelling by blending captured and generative media. The system supports creators through context-aware generation, narrative scaffolding, and creative controls, enabling them to fill narrative gaps and explore new storytelling possibilities. In a user study, Vidmento augmented creators' workflows, fostering creativity while maintaining narrative coherence. Despite challenges in refining outputs and blending styles, the tool demonstrates the potential of AI-assisted video authoring to expand creative expression. Future work will focus on enhancing adaptability, exploring new blending dimensions, and refining user interfaces for hybrid storytelling.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Rewriting Video: Text-Driven Reauthoring of Video Footage
IUI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
VideoDiff: Human-AI Video Co-Creation with Alternatives
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 71%
Protosampling: Enabling Free-Form Convergence of Sampling and Prototyping through Canvas-Driven Visual AI Generation
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
Videostrates: Collaborative, Distributed and Programmable Video Manipulation
UIST '19· Video Production & Editing +1
- 67%
Understanding the Dynamics in Deploying AI-Based Content Creation Support Tools in Broadcasting Systems - Benefits, Challenges, and Directions
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 67%
Videogenic: Identifying Highlight Moments in Videos with Professional Photographs as a Prior
C&C '24· Generative AI (Text, Image, Music, Video) +1
- 67%
Generative Rotoscoping: A First-Person Autobiographical Exploration on Generative Video-to-Video Practices
C&C '25· Generative AI (Text, Image, Music, Video) +1
- 63%
VidTune: Creating Video Soundtracks with Generative Music and Video-Based Thumbnails
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 63%
“It’s more of a vibe I’m going for”: Designing Text-to-Music Generation Interfaces for Video Creators
DIS '25· Generative AI (Text, Image, Music, Video) +3
Based on Jaccard similarity of research subtopics & professions (≥60%)