Do Entropic Measurements of the Diversity of AI-generated Images Match Human Judgement?
Authors
Paper Title
Do Entropic Measurements of the Diversity of AI-generated Images Match Human Judgement?
Publication Info
- Topic area: Measuring and understanding diversity in AI-generated images for creative applications.
- Keywords: Text-to-image models, diversity measurement, entropy-based metrics, human judgment, creative tasks, Stable Diffusion, generative AI, HCI, algorithmic diversity, thematic analysis.
Background and Problem
- Problem / challenge: Existing benchmarks for text-to-image models focus primarily on image quality and fit-to-prompt, neglecting the diversity of generated outputs. Current diversity measures often require large reference datasets or many samples, making them unsuitable for interactive applications. Moreover, these measures have not been sufficiently validated against human judgments.
- Significance: Diversity is critical for supporting creativity in open-ended tasks, where users benefit from a broad range of options to explore novel ideas. Without robust diversity measures, generative AI systems may fail to adequately support ideation and divergent thinking.
- Motivation and related work: Prior work has explored novelty and diversity in generative AI but lacks robust standards or validated measures. Existing metrics like Inception Score and Frechét Inception Distance focus on dataset coverage rather than diversity in responses to specific prompts. This paper builds on entropy-based diversity measures and aims to validate their alignment with human judgments.
Solution
- Proposed approach: The paper evaluates entropy-based diversity measures (Truncated Entropy, Rényi Kernel Entropy, and Vendi Score) for small sets of AI-generated images and compares them to human diversity judgments.
- Novelty:
- Development of a noise-controlled dataset for benchmarking image diversity.
- Validation of entropy-based diversity measures against human judgments.
- Identification of six key themes underlying human diversity judgments through qualitative analysis.
- Insights into the limitations and applicability of diversity measures in creative tasks.
- Procedure and key techniques:
- Generate image sets with varying noise levels using Stable Diffusion.
- Apply entropy-based diversity measures (TE, RKE, VS) to evaluate diversity.
- Conduct three user studies: quantitative ranking of diversity (Studies 1 and 2) and qualitative analysis of human diversity judgments (Study 3).
- Analyze correlations between algorithmic measures, noise levels, and human rankings using Spearman’s ρ and Kendall’s τ.
- Perform thematic analysis of participant descriptions of diversity.
Results
- Concrete findings:
- Study 1: Strong correlation between human rankings and algorithmic measures for large differences in diversity (Spearman’s ρ = 0.896, Kendall’s τ = 0.882).
- Study 2: Moderate alignment for fine-grained diversity differences, with humans and algorithms failing on the same prompts (Spearman’s ρ = 0.225, Kendall’s τ = 0.201).
- Study 3: Six themes identified in human diversity judgments: focus, details, context, composition, style, and vibes/affect.
- Advantage over baselines: Entropy-based measures (TE, VS) effectively approximate human diversity judgments, even for small image sets. RKE is less practical due to high sample requirements but offers differentiability for training models.
- Experiments / evaluation:
- Study 1: 31 participants ranked image sets with large noise differences.
- Study 2: 162 participants ranked sets with subtle noise differences.
- Study 3: 80 participants described similarities and differences in image sets, analyzed using thematic analysis.
- Metrics: Spearman’s ρ, Kendall’s τ, and thematic coding.
- Limitations and future work:
- Exclusion of human subjects in images limits generalizability.
- Noise-based diversity generation may not generalize across prompts or models.
- Lack of confirmatory factor analysis for identified diversity themes.
- Future work should explore diversity in images with human subjects and test whether diversity enhances creativity.
Summary
This paper highlights the importance of diversity as a benchmark for text-to-image models in creative tasks. It validates entropy-based diversity measures (TE, VS) against human judgments using a noise-controlled dataset and identifies six themes underlying human diversity judgments. While these measures align well with human perceptions, especially for large diversity differences, challenges remain for fine-grained distinctions and prompts with unstable outputs. The findings provide a foundation for improving diversity metrics and designing generative AI systems that better support exploration and ideation.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Re-examining Whether, Why, and How Human-AI Interaction Is Uniquely Difficult to Design
CHI '20· Generative AI (Text, Image, Music, Video) +2
- 71%
“I’m happy even though it’s not real”: GenAI Photo Editing as a Remembering Experience
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
Preference-Guided Prompt Optimization for Text-to-Image Generation
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
Less Redraw, More Explore: Suggestion and Completion for Sketch-to-Image
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
Partnering with Generative AI: Experimental Evaluation of Model-Led and Human-Led Interaction in Human-AI Co-Creation
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas
IUI '26· Generative AI (Text, Image, Music, Video) +2
- 67%
AdaptiveSliders: User-aligned Semantic Slider-based Editing of Text-to-Image Model Output
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 67%
VizCrit: Exploring Strategies for Displaying Computational Feedback in a Visual Design Tool
CHI '26· Generative AI (Text, Image, Music, Video) +1
- 67%
Building Human–Multi-Agent Teams for Creative Works
CHI '26· Generative AI (Text, Image, Music, Video) +1
- 67%
Beyond Productivity: Rethinking the Impact of Creativity Support Tools
C&C '25· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)