Collaposer: Transforming Photo Collections into Visual Assets for Storytelling with Collages

Graphic Design & Typography ToolsCreative Collaboration & Feedback SystemsAI-Assisted Writing & Text GenerationUI/UX DesignersVisual Artists & DesignersContent Creators (YouTubers, Podcasters)

Paper Title

Collaposer: Transforming Photo Collections into Visual Assets for Storytelling with Collages

Publication Info

  • Topic area: Digital tools for visual storytelling and collage creation
  • Keywords: digital collage, storytelling, visual assets, instance segmentation, large language models, semantic clustering, creative tools, asset preparation, animation, user study

Background and Problem

  • Problem / challenge: Preparing visual assets for collage-based storytelling is labor-intensive, involving inefficient photo search, manual image cutouts, and complex organization of large collections.
  • Significance: Streamlining asset preparation can free creators to focus on storytelling and composition, enhancing creative workflows in art, media, and design.
  • Motivation and related work: Existing tools support whole-image manipulation or retrieval but lack fine-grained asset preparation for collage creation. Prior research on photo summarization, design material extraction, and visual asset management has not addressed the specific needs of collage creators, leaving a gap in automating object-level asset preparation.

Solution

  • Proposed approach: Collaposer, a tool that transforms photo collections into organized, ready-to-use visual cutouts based on user-provided story descriptions.
  • Novelty:
    1. A pipeline combining instance segmentation, semantic clustering, and LLM-based association to automate asset preparation.
    2. A hierarchical organization of visual assets into categories (characters, backgrounds, accessories) aligned with story descriptions.
    3. A user interface supporting efficient navigation, selection, and composition of curated assets.
  • Procedure and key techniques:
    • Stage I: Preprocess photo collections by tagging, detecting, and segmenting visual elements using RAM, Grounding DINO, and SAM models.
    • Stage II: Select relevant elements with GPT-4o, classify them into semantic categories, and cluster them hierarchically.
    • Stage III: Present assets in a tree view and canvas, with sizes reflecting selection scores based on diversity, story consistency, and resolution.
    • Export structured outputs (e.g., JSON files, layered images) for further editing in tools like Adobe After Effects.

Results

  • Concrete findings:
    • Collaposer achieved 100% one-pass success in user tests, with participants completing tasks without re-prompting.
    • Participants rated Collaposer higher than baselines in selection consistency (Q1-1: 7.0/7.0), diversity (Q2-1: 6.8/7.0), and usability (Q4-1: 6.9/7.0).
  • Advantage over baselines:
    • Outperformed Ablated-Select in story alignment and diversity by supplementing user prompts with inferred elements.
    • Outperformed Ablated-Present in presentation effectiveness by clustering and resizing assets for better navigation.
  • Experiments / evaluation:
    • User study (N = 12) comparing Collaposer with two ablated baselines across 36 collage stories.
    • Metrics: prompt attempts, task duration, Likert-scale ratings, and qualitative feedback.
    • Participants included 9 amateurs and 3 professionals, with diverse creative backgrounds.
  • Limitations and future work:
    • Misalignment between prompts and user intent in some cases.
    • Lack of transparency in how assets are selected.
    • Single-pass preparation limits iterative refinement.
    • Short study duration limits understanding of long-term engagement.

Summary

Collaposer is a novel tool that automates the preparation of visual assets for collage-based storytelling by leveraging instance segmentation, semantic clustering, and LLM-based reasoning. It addresses inefficiencies in traditional workflows, enabling users to quickly obtain and organize story-aligned cutouts. A user study demonstrated its effectiveness in improving selection consistency, diversity, and usability compared to baselines. While limitations remain in prompt alignment and iterative refinement, Collaposer shows promise for applications in art, media, and education, supporting both static and animated collages.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222281/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791160
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Graphic Design & Typography Tools, Creative Collaboration & Feedback Systems, AI-Assisted Writing & Text Generation
work
Professions
UI/UX Designers, Visual Artists & Designers, Content Creators (YouTubers, Podcasters)
article
Content Status
Full text indexed
hub
Related Papers
1 related papers