Script2Screen: Supporting Dialogue-Centric Scriptwriting with Interactive Audiovisual GenerationScriptwriting has traditionally been text-centric, a modality that only partially conveys the produced audiovisual experience. A formative study with professional writers informed us that connecting textual and audiovisual modalities can aid ideation and iteration, especially for writing dialogues. In this work, we pr…2026ZWZhecheng Wang et al.University of TorontoAI-Assisted Creative WritingVideo Production & EditingCreative Collaboration & Feedback SystemsIUI
VidTune: Creating Video Soundtracks with Generative Music and Video-Based ThumbnailsMusic shapes the tone of videos, yet creators find it hard to find soundtracks that match their video's mood and narrative. Recent text-to-music models let creators generate music from text prompts, but our formative study (N=8) shows creators struggle to construct diverse prompts, quickly review and compare tracks, a…2026MHMina Huh et al.University of Texas, AustinGenerative AI (Text, Image, Music, Video)Music Composition & Sound Design ToolsVideo Production & EditingCHI
Vidmento: Scaffolded Expansion for Video Storytelling with Generative VideoVideo storytelling is often constrained by available material, limiting creative expression and leaving undesired narrative gaps. Generative video offers a new way to address these limitations by augmenting captured media with tailored visuals. To explore this potential, we interviewed eight video creators to identify…2026CYCatherine Yeh et al.Harvard UniversityGenerative AI (Text, Image, Music, Video)Video Production & EditingCreative Collaboration & Feedback SystemsCHI
MapStory: Prototyping Editable Map Animations with LLM AgentsWe introduce MapStory, an LLM‑powered animation prototyping tool that generates editable map animation sequences directly from natural language text by leveraging a dual-agent LLM architecture. Given a user-written script, MapStory automatically produces a scene breakdown, which decomposes the text into key map animat…2025AGAditya Gunturu et al.University of CalgaryGeospatial & Map VisualizationComputational Methods in HCIUIST
SimTube: Simulating Audience Feedback on Videos using Generative AI and User PersonasAudience feedback is crucial for refining video content, yet it typically comes after publication, limiting creators' ability to make timely adjustments. To bridge this gap, we introduce SimTube, a generative AI system designed to simulate audience feedback in the form of video comments before a video's release. SimTu…2025YHYu-Kai Hung et al.National Taiwan UniversityGenerative AI (Text, Image, Music, Video)Live Streaming & Content CreatorsAI-Assisted Creative WritingIUI
GazeNoter: Co-Piloted AR Note-Taking via Gaze Selection of LLM Suggestions to Match Users' IntentionsNote-taking is critical during speeches and discussions, serving for later summarization and organization and for real-time question and opinion reminding in question-and-answer sessions or timely contributions in discussions. Manually typing on smartphones for note-taking could be distracting and increase cognitive l…2025HTHsin-Ruey Tsai et al.National Chengchi UniversityEye Tracking & Gaze InteractionMixed Reality WorkspacesHuman-LLM CollaborationCHI
LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video EditingVideo creation has become increasingly popular, yet the expertise and effort required for editing often pose barriers to beginners. In this paper, we explore the integration of large language models (LLMs) into the video editing workflow to reduce these barriers. Our design vision is embodied in LAVE, a novel system t…2024BWBryan Wang et al.University of TorontoHuman-LLM CollaborationVideo Production & EditingIUI
SynthScribe: Deep Multimodal Tools for Synthesizer Sound Retrieval and ExplorationSynthesizers are powerful tools that allow musicians to create dynamic and original sounds. Existing commercial interfaces for synthesizers typically require musicians to interact with complex low-level parameters or to manage large libraries of premade sounds. To address these challenges, we implement SynthScribe ---…2024SBStephen Brade et al.University of TorontoGenerative AI (Text, Image, Music, Video)Music Composition & Sound Design ToolsCreative Collaboration & Feedback SystemsIUI
Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language ModelsText-to-image generative models have demonstrated remarkable capabilities in generating high-quality images based on textual prompts. However, crafting prompts that accurately capture the user's creative intent remains challenging. It often involves laborious trial-and-error procedures to ensure that the model interpr…2023SBStephen Brade et al.University of TorontoGenerative AI (Text, Image, Music, Video)Human-LLM CollaborationAI-Assisted Creative WritingUIST
Stargazer: An Interactive Camera Robot for Capturing How-To Videos Based on Subtle Instructor CuesLive and pre-recorded video tutorials are an effective means for teaching physical skills such as cooking or prototyping electronics. A dedicated cameraperson following an instructor’s activities can improve production quality. However, instructors who do not have access to a cameraperson’s help often have to work wi…2023JLJiannan Li et al.Singapore Management UniversityTeleoperation & TelepresenceCHI
Enabling Conversational Interaction with Mobile UI using Large Language ModelsConversational agents show the promise to allow users to interact with mobile devices using language. However, to perform diverse UI tasks with natural language, developers typically need to create separate datasets and models for each specific task, which is expensive and effort-consuming. Recently, pre-trained large…2023BWBryan Wang et al.University of TorontoVoice User Interface (VUI) DesignHuman-LLM CollaborationCHI
Creepy Assistant: Development and Validation of a Scale to Measure the Perceived Creepiness of Voice AssistantsVoice assistants have afforded users rich interaction opportunities to access information and issue commands in a variety of contexts. However, some users feel uneasy or creeped out by voice assistants, leading to a decreased desire to use them. As there has yet to be a comprehensive understanding of the factors that…2023RPRachel Phinnemore et al.University of TorontoVoice User Interface (VUI) DesignAgent Personality & AnthropomorphismExplainable AI (XAI)CHI
Record Once, Post Everywhere: Automatic Shortening of Audio Stories for Social MediaFollowing the prevalence of short-form video, short-form voice content has emerged on social media platforms like Twitter and Facebook. A challenge that creators face is hard constraints on the content length. If the initial recording is not short enough, they need to re-record or edit their content. Both are time-con…2022BWBryan Wang et al.University of TorontoVoice User Interface (VUI) DesignConversational ChatbotsAI-Assisted Decision-Making & AutomationUIST
Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningMobile User Interface Summarization generates succinct language descriptions of mobile screens for conveying important contents and functionalities of the screen, which can be useful for many language-based application scenarios. We present Screen2Words, a novel screen summarization approach that automatically encapsu…2021BWBryan Wang et al.University of TorontoVoice User Interface (VUI) DesignHuman-LLM CollaborationUIST
Soloist: Generating Mixed-Initiative Tutorials from Existing Guitar Instructional Videos Through Audio ProcessingLearning musical instruments using online instructional videos has become increasingly prevalent. However, pre-recorded videos lack the instantaneous feedback and personal tailoring that human tutors provide. In addition, existing video navigations are not optimized for instrument learning, making the learning experie…2021BWBryan Wang et al.University of TorontoFitness Tracking & Physical Activity MonitoringMusic Composition & Sound Design ToolsCreative Collaboration & Feedback SystemsCHI
BlyncSync: Enabling Multimodal Smartwatch Gestures with Synchronous Touch and BlinkInput techniques have been drawing abiding attention along with the continual miniaturization of personal computers. In this paper, we present BlyncSync, a novel multi-modal gesture set that leverages the synchronicity of touch and blink events to augment the input vocabulary of smartwatches with a rapid gesture, whil…2020BWBryan Wang et al.University of TorontoSmartwatches & Fitness BandsCHI
ActiveErgo: Automatic and Personalized Ergonomics using Self-actuating FurnitureProper ergonomics improves productivity and reduces risks for injuries such as tendinosis, tension neck syndrome, and back injuries. Despite having ergonomics standards and guidelines for computer usage since the 1980s, injuries due to poor ergonomics remain widespread. We present ActiveErgo, the first active approac…2018YWYu-Chian Wu et al.National Taiwan UniversityFull-Body Interaction & Embodied InputKnowledge Worker Tools & WorkflowsCHI