Do Entropic Measurements of the Diversity of AI-generated Images Match Human Judgement?

Generative AI (Text, Image, Music, Video)Explainable AI (XAI)Creative Collaboration & Feedback SystemsAI/ML Researchers & EngineersHCI ResearchersUI/UX Designers

Paper Title

Do Entropic Measurements of the Diversity of AI-generated Images Match Human Judgement?

Publication Info

  • Topic area: Measuring and understanding diversity in AI-generated images for creative applications.
  • Keywords: Text-to-image models, diversity measurement, entropy-based metrics, human judgment, creative tasks, Stable Diffusion, generative AI, HCI, algorithmic diversity, thematic analysis.

Background and Problem

  • Problem / challenge: Existing benchmarks for text-to-image models focus primarily on image quality and fit-to-prompt, neglecting the diversity of generated outputs. Current diversity measures often require large reference datasets or many samples, making them unsuitable for interactive applications. Moreover, these measures have not been sufficiently validated against human judgments.
  • Significance: Diversity is critical for supporting creativity in open-ended tasks, where users benefit from a broad range of options to explore novel ideas. Without robust diversity measures, generative AI systems may fail to adequately support ideation and divergent thinking.
  • Motivation and related work: Prior work has explored novelty and diversity in generative AI but lacks robust standards or validated measures. Existing metrics like Inception Score and Frechét Inception Distance focus on dataset coverage rather than diversity in responses to specific prompts. This paper builds on entropy-based diversity measures and aims to validate their alignment with human judgments.

Solution

  • Proposed approach: The paper evaluates entropy-based diversity measures (Truncated Entropy, Rényi Kernel Entropy, and Vendi Score) for small sets of AI-generated images and compares them to human diversity judgments.
  • Novelty:
    1. Development of a noise-controlled dataset for benchmarking image diversity.
    2. Validation of entropy-based diversity measures against human judgments.
    3. Identification of six key themes underlying human diversity judgments through qualitative analysis.
    4. Insights into the limitations and applicability of diversity measures in creative tasks.
  • Procedure and key techniques:
    1. Generate image sets with varying noise levels using Stable Diffusion.
    2. Apply entropy-based diversity measures (TE, RKE, VS) to evaluate diversity.
    3. Conduct three user studies: quantitative ranking of diversity (Studies 1 and 2) and qualitative analysis of human diversity judgments (Study 3).
    4. Analyze correlations between algorithmic measures, noise levels, and human rankings using Spearman’s ρ and Kendall’s τ.
    5. Perform thematic analysis of participant descriptions of diversity.

Results

  • Concrete findings:
    • Study 1: Strong correlation between human rankings and algorithmic measures for large differences in diversity (Spearman’s ρ = 0.896, Kendall’s τ = 0.882).
    • Study 2: Moderate alignment for fine-grained diversity differences, with humans and algorithms failing on the same prompts (Spearman’s ρ = 0.225, Kendall’s τ = 0.201).
    • Study 3: Six themes identified in human diversity judgments: focus, details, context, composition, style, and vibes/affect.
  • Advantage over baselines: Entropy-based measures (TE, VS) effectively approximate human diversity judgments, even for small image sets. RKE is less practical due to high sample requirements but offers differentiability for training models.
  • Experiments / evaluation:
    • Study 1: 31 participants ranked image sets with large noise differences.
    • Study 2: 162 participants ranked sets with subtle noise differences.
    • Study 3: 80 participants described similarities and differences in image sets, analyzed using thematic analysis.
    • Metrics: Spearman’s ρ, Kendall’s τ, and thematic coding.
  • Limitations and future work:
    • Exclusion of human subjects in images limits generalizability.
    • Noise-based diversity generation may not generalize across prompts or models.
    • Lack of confirmatory factor analysis for identified diversity themes.
    • Future work should explore diversity in images with human subjects and test whether diversity enhances creativity.

Summary

This paper highlights the importance of diversity as a benchmark for text-to-image models in creative tasks. It validates entropy-based diversity measures (TE, VS) against human judgments using a noise-controlled dataset and identifies six themes underlying human diversity judgments. While these measures align well with human perceptions, especially for large diversity differences, challenges remain for fine-grained distinctions and prompts with unstable outputs. The findings provide a foundation for improving diversity metrics and designing generative AI systems that better support exploration and ideation.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223505/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791383
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Explainable AI (XAI), Creative Collaboration & Feedback Systems
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers