From Camera-Eye to AI: Exploring the Interplay of Cinematography and Computational Visual Storytelling

Generative AI (Text, Image, Music, Video)AI-Assisted Creative WritingVideo Production & EditingFilm & Animation ProducersPodcast ProducersVisual Artists & Designers

Research Background and Problem

  • What problems or challenges did the authors identify?
    Existing computational visual narratives primarily focus on analyzing image content while neglecting the role of formal elements such as composition, camera movement, and lighting. The authors raise a critical question: How do cinematographic techniques influence AI's interpretation of images and the narratives it generates?

  • Why is this issue important?
    Formal elements are central to cinematic storytelling, as they not only convey narrative atmosphere but also deepen the understanding of content. These perceptual layers could significantly impact AI-generated visual narratives. However, this area has not received sufficient attention to date. Combining cinematography with AI narratives could provide new insights for film studies and computational vision system design.

  • Research Motivation and Related Work

    • Core Motivation: To investigate how the formal characteristics of cinematographic techniques shape AI's performance in visual and narrative dimensions, complementing existing content-based research.
    • The authors connect the interdisciplinary fields of visual storytelling, film studies, and HCI (Human-Computer Interaction), proposing that visual narrative systems should not only focus on content generation but also enhance the exploration of the complexity between form and meaning.

Solution

  • What methods or solutions did the authors propose?
    Using 60 static image frames from the film Man with a Movie Camera (1929), the authors input multidimensional prompts into a visual language model (VLM) to generate AI interpretations and narratives. This approach explores three main themes:

    1. How AI perceives "drama" and "power" in social realities through camera angles and types.
    2. How AI interprets and (mis)understands the ambiguity or clarity of reality brought about by lighting and focus.
    3. How AI processes and generates surreal narratives through visual effects (e.g., double exposure).
  • What is innovative about this solution?

    1. The authors combine cinematographic research methods (e.g., shot analysis) with AI visual storytelling, opening a new interdisciplinary research direction.
    2. They propose a "contextualized gaze" approach, systematically exploring the role of formal elements in shaping AI's understanding and narrative generation by linking AI-generated text back to cinematographic techniques and their narrative impact.
    3. They employ a combination of "close reading" and thematic analysis to deeply analyze the multi-layered meanings of each AI-generated story, providing a new methodological paradigm that integrates literary and technical approaches.
  • What are the implementation steps and key technologies used?

    1. Image Selection: 60 representative frames showcasing visual formal techniques were selected from Man with a Movie Camera.
    2. Prompt Engineering: A five-step prompt design (including observation, contextualization, planning, narration, and titling) was created to guide the VLM in generating multidimensional outputs.
    3. Data Generation and Analysis: The Claude 3.5 AI model was used to input images and prompts, generating narratives that were then deconstructed using "close reading" and thematic analysis methods.
    4. Comparative Experiments: A/B testing (e.g., varying lighting and focus) was conducted to distinguish the impact of these techniques on AI narrative generation.

Research Findings

  • What specific findings were achieved?

    • The study revealed how AI-generated stories are shaped by various formal techniques such as camera angles, focus, and visual effects.
    • Three main themes were identified:
      1. The relationship between camera perspectives and the perception of dramatic conflict and power in social realities.
      2. How lighting and focus influence AI's interpretation of ambiguity and clarity in representing reality.
      3. How visual effects (e.g., double exposure) inspire AI to generate multi-layered, surreal narratives.
  • What advantages does this solution have compared to existing ones?

    1. Form-Oriented Approach: Addresses the gap in existing visual storytelling research by emphasizing formal elements.
    2. Multidimensional Perspective: Deepens the understanding of the interplay between "form and content" in AI interpretation and narrative generation.
    3. Contextualized Perception: Introduces a contextual analysis framework that integrates cinematography with narrative generation.
    4. Integration of Art and Narrative: Expands AI applications in art and storytelling through the analysis of surreal and multi-layered narratives.
  • What were the experimental or evaluation results?
    The experiments demonstrated that:

    • Camera angles significantly influence AI's interpretation of power dynamics (e.g., high/low-angle perspectives) and emotional intensity.
    • Images with ambiguous lighting and focus can lead to AI generating narratives that deviate from the actual content, while also showcasing AI's creative ability to handle ambiguity and abstract elements.
    • The multi-layered expression of visual effects (e.g., double exposure) can guide AI in generating surreal narratives, presenting complex imagery and symbolism through layered text.
  • Limitations and Future Directions

    • Limitations:

      1. The study focuses on static frames from a single film, leaving the narrative potential of dynamic imagery unexplored.
      2. Some AI-generated content exhibits subjectivity, which may lead to discrepancies with user interpretations.
      3. The generalizability of the extracted patterns needs to be validated across more films or image datasets.
    • Future Directions:

      1. Develop AI visual storytelling platforms that support dynamic imagery and temporal sequences.
      2. Create user-interactive cinematographic control tools, allowing users to simulate cinematographic parameters to explore different narrative outputs.
      3. Investigate how cinematographic techniques influence other AI tasks, such as narrative generation for journalism, education, or archiving.

Conclusion

This study combines cinematography with visual language models to reveal how formal elements profoundly influence AI narrative generation. It not only provides a theoretical foundation for the design of future visual storytelling systems but also demonstrates the potential of AI to reimagine history through surreal narratives.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188294/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713840
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), AI-Assisted Creative Writing, Video Production & Editing
work
Professions
Film & Animation Producers, Podcast Producers, Visual Artists & Designers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers