An Exploration of Default Images in Text-to-Image Generation

Honorable Mention
Generative AI (Text, Image, Music, Video)Human-LLM CollaborationAI-Assisted Decision-Making & AutomationExplainable AI (XAI)AI/ML Researchers & EngineersUI/UX DesignersHCI Researchers

Paper Title

An Exploration of Default Images in Text-to-Image Generation

Publication Info

  • Topic area: Investigation of default image phenomena in text-to-image (TTI) generation systems.
  • Keywords: Text-to-image generation, default images, Midjourney, prompt engineering, generative AI, user satisfaction, computational creativity, latent space, mode collapse, human-AI interaction.

Background and Problem

  • Problem / challenge: TTI models generate default images—visually similar outputs from unrelated prompts—when encountering unknown or ambiguous terms. This phenomenon is underexplored and impacts user satisfaction and creative workflows.
  • Significance: Understanding default images can improve prompt engineering, model robustness, and user experience, while addressing gaps in computational creativity and semantic alignment.
  • Motivation and related work: Prior research has focused on prompt engineering, failure modes, and user satisfaction in TTI systems, but default images remain unexamined. This paper builds on studies of mode collapse and generative model biases to explore default images systematically.

Solution

  • Proposed approach: A systematic investigation of default images in Midjourney, combining manual experiments, user studies, and large-scale computational analysis.
  • Novelty:
    1. Definition and empirical characterization of default images, including six prompt categories likely to trigger them.
    2. Development of a scalable method to identify default images using clustering and filtering techniques.
    3. Creation of a dataset with 189,432 prompts and 757,728 images for future research.
    4. Identification of eight postulates describing default image characteristics.
  • Procedure and key techniques:
    1. Manual creation of 130 prompts across six categories (e.g., rare names, low-resource languages, glitch tokens).
    2. Image generation experiments with fixed parameters and ablation studies to test variability.
    3. Affinity diagramming to identify default images based on visual similarity.
    4. Online user study (N=48) to assess the impact of default images on user satisfaction.
    5. Computational analysis of 757,728 images from Midjourney to identify default image clusters using CLIP embeddings, clustering, and filtering.

Results

  • Concrete findings:
    • Identified ten canonical default images in manual experiments, recurring across unrelated prompts.
    • Computational analysis revealed 4,715 clusters of default images (36,243 images, 4.84% of the dataset).
    • Default images often feature human-animal or human-vegetal fusions, dream-like motifs, and aesthetic biases.
    • User study showed significant dissatisfaction with default images (e.g., Q4 satisfaction mean = 2.4, SD = 1.9).
  • Advantage over baselines:
    • First systematic study of default images, providing empirical evidence and scalable detection methods.
    • Demonstrated real-world relevance of default images in creative user prompts.
  • Experiments / evaluation:
    • Manual experiments with 130 prompts, generating 520 images.
    • Ablation studies on seed values, style modifiers, and model versions.
    • User study measuring satisfaction with default images (N=48).
    • Computational analysis of 757,728 images using clustering and semantic filtering.
  • Limitations and future work:
    • Limited to Midjourney; findings may not generalize to other TTI systems.
    • Lack of access to model internals and training data restricts causal interpretation.
    • Future work could explore default images in other systems, personalization effects, and real-time detection mechanisms.

Summary

This paper presents the first systematic investigation of default images in Midjourney, a TTI generation system. Default images, which arise from ambiguous or unknown prompts, were characterized through manual experiments, user studies, and computational analysis of 757,728 images. The study identified recurring motifs, user dissatisfaction with default images, and eight postulates describing their properties. By defining and analyzing default images, the paper highlights their impact on user experience and creative workflows, providing a foundation for improving generative AI systems and supporting future research.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222956/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790681
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
5 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, Explainable AI (XAI)
work
Professions
AI/ML Researchers & Engineers, UI/UX Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers