Characterizing Photorealism and Artifacts in Diffusion Model-Generated Images

Generative AI (Text, Image, Music, Video)Explainable AI (XAI)Deepfake & Synthetic Media DetectionData Scientists & AnalystsAI/ML Researchers & EngineersHCI Researchers

Research Background and Problems

  • What issues or challenges did the authors identify?
    Images generated by diffusion models are becoming increasingly realistic, often indistinguishable from real ones. However, these AI-generated images typically contain artifacts and inconsistencies, such as anatomically implausible features, lighting errors, or physically unrealistic scenes. These flaws provide potential clues for detection by humans and machines, but it remains unclear how these issues manifest across different scenarios and how human ability to identify them varies.

  • Why is this issue important?
    With the widespread application of generative AI, the number of forged images is increasing, potentially undermining public trust in image authenticity and media credibility. Highly realistic fake images could be used to spread misinformation, leading to negative societal impacts.

  • Research Motivation and Related Work
    Existing detection methods, such as machine learning models, often struggle with robustness when AI-generated images undergo post-processing modifications (e.g., cropping or compression). Furthermore, while there is extensive research on GAN-generated images, studies on diffusion models remain relatively limited. The motivation of this research is to uncover the capabilities and limitations of diffusion image generation models through human experiments and to develop an interpretable artifact classification method to help the public identify AI-generated images.

Solutions

  • What methods or solutions did the authors propose?

    1. Proposed a human perception-based measurement of "image realism" (objectively quantified through human accuracy in distinguishing real from fake images).
    2. Developed an artifact classification method to systematically categorize visible features in images generated by diffusion models.
    3. Conducted a large-scale online experiment, collecting classification results from 50,444 participants for 599 images (including AI-generated and real photos) to evaluate human detection capabilities for AI-generated images.
  • What are the innovative aspects of this solution?

    1. Introduced a novel artifact classification method encompassing multidimensional features such as anatomical inconsistencies, violations of physical laws, and socio-cultural implausibilities.
    2. Quantified the abstract concept of "realism" in the framework of human psychophysics by using accuracy as an inverse measure, effectively avoiding biases associated with subjective questions (e.g., "Does this image look real?").
    3. For the first time, employed large-scale experiments to analyze various factors influencing human judgment, such as scene complexity, display time, and manual curation's impact on model realism.
  • What are the implementation steps and key technologies used?

    1. Development of the artifact classification method: Conducted preliminary literature review and generated thousands of AI images to summarize five major artifact categories: anatomical inconsistencies, stylistic artifacts, functional implausibilities, physical violations, and socio-cultural implausibilities.
    2. Image generation and curation: Generated 450 images using Midjourney, Firefly, and Stable Diffusion, combined with 149 real photos to construct the experimental dataset.
    3. Experiment design and data collection: Developed an online experimental platform, employing randomized image display times and collecting participants' accuracy in real-time.
    4. Data analysis: Conducted quantitative and qualitative analyses to compare detection accuracy across different artifact categories and examined trends in human judgment capabilities with varying display times.

Research Findings

  • What specific findings were achieved?

    1. Classification accuracy statistics showed that participants achieved an average detection accuracy of 76% for AI-generated images and 74% for real images. However, accuracy significantly decreased in complex scenes (e.g., group portraits).
    2. Human ability to recognize artifacts was strongly correlated with artifact categories; anatomical errors (e.g., finger or facial details) and stylistic artifacts (e.g., plastic textures) were easiest to detect, while functional errors (e.g., slack guitar strings) and physical violations were more challenging.
    3. Display time significantly influenced human accuracy: accuracy was 72% for 1-second display times, increasing to 82% when display time was unrestricted.
    4. Manually curated high-quality images were harder to identify as fake compared to randomly generated ones, highlighting the role of human intervention in enhancing image realism.
  • What advantages does this solution have compared to existing ones?

    1. Provides an interpretable artifact classification method with high generalizability and applicability to visual errors perceived by humans.
    2. Validates the effectiveness and scalability of the classification method through large-scale human experiments and anonymous participant feedback.
    3. Reveals how real-world parameters like scene complexity and display time influence the detection of AI-generated images, offering practical guidance for application.
  • What are the experimental or evaluation results?

    • AI images in complex scenes (e.g., groups or dynamic backgrounds) are easier to identify.
    • Display times longer than 5 seconds significantly improve detection accuracy.
    • Manual curation can partially mitigate the limitations of generative AI image quality.
  • Limitations and Future Directions

    1. Limitations
      • The study's participants were self-selected and widely distributed, potentially introducing cultural or life background biases, as demographic information was not deeply collected.
      • The research focused on mainstream generative AI models from 2024 (Midjourney, Firefly, and Stable Diffusion); future studies may need to validate findings with next-generation models and associated artifacts.
      • The study primarily addressed image recognition, leaving video and audio content forgery unexplored.
    2. Future Directions
      • Explore how the artifact classification method can be integrated into AI detection tools or AI literacy education modules.
      • Investigate the applicability of the classification method to videos and dynamic scenes, examining the characteristics of temporal sequence artifacts across frames.
      • Design intervention experiments to test whether education based on the classification method can enhance human ability to identify AI-generated content.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188583/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713962
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Explainable AI (XAI), Deepfake & Synthetic Media Detection
work
Professions
Data Scientists & Analysts, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers