Generative Rotoscoping: A First-Person Autobiographical Exploration on Generative Video-to-Video Practices
Authors
This paper contributes a first-person exploration on AI video-to-video technologies, which I call "Generative Rotoscoping". This includes: insights on the creation process, a set of prototype explorations, and an integrated workflow for multi-modal video generation. Generative video is rapidly evolving and delivering higher quality outputs. While video generation models have potential for film-making and content creation, they lack controllability for creative expression: viable videos can require hundreds of unsuccessful attempts. To understand this emergent practice, and due to the constant evolution of models and limited number of early adopters, I explored Generative Rotoscoping over 12 months and created AI workflows leading to over 40,000 video/image files examining a variety of models and techniques including: structural guidance, frame consistency, image referencing and masks, compositing, among others. Insights from this work can serve as a starting point for designing the next generation of video authoring tools.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 75%
Generating Highlight Videos of a User-Specified Length using Most Replayed Data
CHI '25· Video Production & Editing
- 67%
Vidmento: Scaffolded Expansion for Video Storytelling with Generative Video
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 67%
Rewriting Video: Text-Driven Reauthoring of Video Footage
IUI '26· Generative AI (Text, Image, Music, Video) +2
- 60%
B-Script: Transcript-based B-roll Video Editing with Recommendations
CHI '19· Recommender System UX +1
- 60%
Understanding the Dynamics in Deploying AI-Based Content Creation Support Tools in Broadcasting Systems - Benefits, Challenges, and Directions
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 60%
XCam: Mixed-Initiative Virtual Cinematography for Live Production of Virtual Reality Experiences
CHI '25· Social & Collaborative VR +1
- 60%
Videogenic: Identifying Highlight Moments in Videos with Professional Photographs as a Prior
C&C '24· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)