ExpressEdit: Video Editing with Natural Language and Sketching
Authors
Informational videos serve as a crucial source for explaining conceptual and procedural knowledge to novices and experts alike. When producing informational videos, editors edit videos by overlaying text/images or trimming footage to enhance the video quality and make it more engaging. However, video editing can be difficult and time-consuming, especially for novice video editors who often struggle with expressing and implementing their editing ideas. To address this challenge, we first explored how multimodality — natural language (NL) and sketching, which are natural modalities humans use for expression—can be utilized to support video editors in expressing video editing ideas. We gathered 176 multimodal expressions of editing commands from 10 video editors, which revealed the patterns of use of NL and sketching in describing edit intents. Based on the findings, we present ExpressEdit, a system that enables editing videos via NL text and sketching on the video frame. Powered by LLM and vision models, the system interprets (1) temporal, (2) spatial, and (3) operational references in an NL command and spatial references from sketching. The system implements the interpreted edits, which then the user can iterate on. An observational study (N=10) showed that ExpressEdit enhanced the ability of novice video editors to express and implement their edit ideas. The system allowed participants to perform edits more efficiently and generate more ideas by generating edits based on user’s multimodal edit commands and supporting iterations on the editing commands. This work offers insights into the design of future multimodal interfaces and AI-based pipelines for video editing.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can multimodal interaction combining natural language and hand-drawn sketches help users more accurately express video editing needs?Category: Multimodal Video Editing Expression and ControlSimilar questionsarrow_forward
- Can the ExpressEdit system meet user expectations for parsing precision in time, space, and editing parameters?Category: Multimodal Video Editing Expression and ControlSimilar questionsarrow_forward
- How much can multimodal interaction reduce video editing complexity and improve editing efficiency?Category: Multimodal Video Editing Expression and ControlSimilar questionsarrow_forward
Practical Problems
1- Beginners in informational video editing often face operational complexity and difficulty expressing intent.Category: Multimodal Video Editing Expression and ControlSimilar questionsarrow_forward
- 67%
Understanding the Dynamics in Deploying AI-Based Content Creation Support Tools in Broadcasting Systems - Benefits, Challenges, and Directions
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 67%
Videogenic: Identifying Highlight Moments in Videos with Professional Photographs as a Prior
C&C '24· Generative AI (Text, Image, Music, Video) +1
- 67%
Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language Models
UIST '23· Generative AI (Text, Image, Music, Video) +2
- 63%
VidSTR: Automatic Spatiotemporal Retargeting of Speech-Driven Video Compositions
CHI '25· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)