Vidmento: Scaffolded Expansion for Video Storytelling with Generative Video

Generative AI (Text, Image, Music, Video)Video Production & EditingCreative Collaboration & Feedback SystemsContent Creators (YouTubers, Podcasters)Film & Animation ProducersUI/UX Designers

Paper Title

Vidmento: Creating Video Stories through Context-Aware Expansion with Generative Video

Publication Info

  • Topic area: Hybrid video storytelling using generative and captured media
  • Keywords: Generative video, hybrid storytelling, video editing, AI-assisted creation, narrative expansion, cinematic principles, creative tools, video authoring, context-aware generation, generative AI

Background and Problem

  • Problem / challenge: Video storytelling is constrained by the availability of captured footage, limiting creators' ability to fill narrative gaps or explore alternative storylines. Existing generative video tools often fail to integrate seamlessly with captured media, lacking contextual blending and narrative coherence.
  • Significance: Addressing these limitations enables creators to craft richer, more cohesive video stories by combining the authenticity of captured footage with the flexibility of generative media.
  • Motivation and related work: Prior tools focus on either fully captured or fully generated workflows, often treating generative video as an isolated process. This paper identifies a gap in hybrid storytelling approaches and aims to bridge it by blending captured and generative media within a unified framework.

Solution

  • Proposed approach: Vidmento, a video authoring tool that implements the "generative expansion" framework to blend captured and generative media for hybrid video storytelling.
  • Novelty:
    1. Introduction of "generative expansion," a design framework for contextually blending captured and generative video.
    2. Development of Vidmento, a tool that integrates narrative and cinematic principles to support hybrid video creation.
    3. Empirical findings from formative interviews and a user study highlighting opportunities and challenges in hybrid storytelling.
  • Procedure and key techniques:
    1. Vidmento provides a semi-structured canvas for organizing and sequencing media, coupled with a script editor for narration.
    2. Context-aware AI generates new video clips and script suggestions to fill narrative gaps or expand stories.
    3. Users retain creative control through features like annotation-based refinements, structured prompts, and timeline alignment tools.
    4. The system employs a two-stage pipeline for generating contextual keyframes and animating them into videos.

Results

  • Concrete findings:
    • Vidmento enabled users to augment their original media by an average of 96.4%, generating ∼31 images and 8 videos per session.
    • The system supported narrative exploration, with creators clicking an average of 3.7 story suggestions and generating 2.0 refined script versions per session.
  • Advantage over baselines:
    • Unlike existing tools, Vidmento integrates captured and generative media within a unified workspace, offering context-aware blending and narrative scaffolding.
    • Users valued its ability to visualize and manipulate stories through a semi-structured canvas and synchronized script editor.
  • Experiments / evaluation:
    • Conducted a user study with 12 creators (novices to professionals) who produced 1–2 minute narrated videos using Vidmento.
    • Participants highlighted the system's utility for ideation, narrative expansion, and blending media, though some noted challenges with refining generative outputs.
  • Limitations and future work:
    • Challenges include style mismatches between captured and generated footage, difficulty refining outputs, and concerns about authenticity in personal stories.
    • Future directions include improving model adaptability, exploring spatial and stylistic blending, and designing fully unified multimodal canvases.

Summary

Vidmento introduces a novel framework, "generative expansion," for hybrid video storytelling by blending captured and generative media. The system supports creators through context-aware generation, narrative scaffolding, and creative controls, enabling them to fill narrative gaps and explore new storytelling possibilities. In a user study, Vidmento augmented creators' workflows, fostering creativity while maintaining narrative coherence. Despite challenges in refining outputs and blending styles, the tool demonstrates the potential of AI-assisted video authoring to expand creative expression. Future work will focus on enhancing adaptability, exploring new blending dimensions, and refining user interfaces for hybrid storytelling.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222152/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791337
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Video Production & Editing, Creative Collaboration & Feedback Systems
work
Professions
Content Creators (YouTubers, Podcasters), Film & Animation Producers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
9 related papers