An Exploration of Default Images in Text-to-Image Generation
Honorable MentionAuthors
Paper Title
An Exploration of Default Images in Text-to-Image Generation
Publication Info
- Topic area: Investigation of default image phenomena in text-to-image (TTI) generation systems.
- Keywords: Text-to-image generation, default images, Midjourney, prompt engineering, generative AI, user satisfaction, computational creativity, latent space, mode collapse, human-AI interaction.
Background and Problem
- Problem / challenge: TTI models generate default images—visually similar outputs from unrelated prompts—when encountering unknown or ambiguous terms. This phenomenon is underexplored and impacts user satisfaction and creative workflows.
- Significance: Understanding default images can improve prompt engineering, model robustness, and user experience, while addressing gaps in computational creativity and semantic alignment.
- Motivation and related work: Prior research has focused on prompt engineering, failure modes, and user satisfaction in TTI systems, but default images remain unexamined. This paper builds on studies of mode collapse and generative model biases to explore default images systematically.
Solution
- Proposed approach: A systematic investigation of default images in Midjourney, combining manual experiments, user studies, and large-scale computational analysis.
- Novelty:
- Definition and empirical characterization of default images, including six prompt categories likely to trigger them.
- Development of a scalable method to identify default images using clustering and filtering techniques.
- Creation of a dataset with 189,432 prompts and 757,728 images for future research.
- Identification of eight postulates describing default image characteristics.
- Procedure and key techniques:
- Manual creation of 130 prompts across six categories (e.g., rare names, low-resource languages, glitch tokens).
- Image generation experiments with fixed parameters and ablation studies to test variability.
- Affinity diagramming to identify default images based on visual similarity.
- Online user study (N=48) to assess the impact of default images on user satisfaction.
- Computational analysis of 757,728 images from Midjourney to identify default image clusters using CLIP embeddings, clustering, and filtering.
Results
- Concrete findings:
- Identified ten canonical default images in manual experiments, recurring across unrelated prompts.
- Computational analysis revealed 4,715 clusters of default images (36,243 images, 4.84% of the dataset).
- Default images often feature human-animal or human-vegetal fusions, dream-like motifs, and aesthetic biases.
- User study showed significant dissatisfaction with default images (e.g., Q4 satisfaction mean = 2.4, SD = 1.9).
- Advantage over baselines:
- First systematic study of default images, providing empirical evidence and scalable detection methods.
- Demonstrated real-world relevance of default images in creative user prompts.
- Experiments / evaluation:
- Manual experiments with 130 prompts, generating 520 images.
- Ablation studies on seed values, style modifiers, and model versions.
- User study measuring satisfaction with default images (N=48).
- Computational analysis of 757,728 images using clustering and semantic filtering.
- Limitations and future work:
- Limited to Midjourney; findings may not generalize to other TTI systems.
- Lack of access to model internals and training data restricts causal interpretation.
- Future work could explore default images in other systems, personalization effects, and real-time detection mechanisms.
Summary
This paper presents the first systematic investigation of default images in Midjourney, a TTI generation system. Default images, which arise from ambiguous or unknown prompts, were characterized through manual experiments, user studies, and computational analysis of 757,728 images. The study identified recurring motifs, user dissatisfaction with default images, and eight postulates describing their properties. By defining and analyzing default images, the paper highlights their impact on user experience and creative workflows, providing a foundation for improving generative AI systems and supporting future research.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 Articles
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 86%
Re-examining Whether, Why, and How Human-AI Interaction Is Uniquely Difficult to Design
CHI '20· Generative AI (Text, Image, Music, Video) +2
- 86%
Effects of LLM-based Search on Decision Making: Speed, Accuracy, and Overreliance
CHI '25· Human-LLM Collaboration +2
- 86%
How Users Perceive Mixed-Initiative AI: Attitudes Toward Assistance in Problem Solving
IUI '26· Human-LLM Collaboration +2
- 86%
Adaptive Prompt Elicitation for Text-to-Image Generation
IUI '26· Generative AI (Text, Image, Music, Video) +2
- 75%
GAM Coach: Towards Interactive and User-centered Algorithmic Recourse
CHI '23· Generative AI (Text, Image, Music, Video) +3
- 75%
Bridging Gulfs in UI Generation through Semantic Guidance
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 75%
PCGEF: A Framework for Diagnosing Subjective Alignment in Human-Centered Persona-Conditioned Generation
CHI '26· Human-LLM Collaboration +3
- 75%
Agentic Audio Moderators vs Humans in Think-Aloud Usability Testing
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 75%
Feedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational Agents
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)