Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI

Explainable AI (XAI)AI-Assisted Decision-Making & AutomationAlgorithmic Transparency & AuditabilityInteractive Data VisualizationAI/ML Researchers & EngineersData Scientists & AnalystsHCI Researchers

Paper Title

Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI

Publication Info

  • Topic area: Systematic auditing of text-to-image generative AI models
  • Keywords: Text-to-image models, AI auditing, scene graph, LLM guidance, systematic auditing, visual analytics, generative AI, human-AI collaboration, bias detection, responsible AI

Background and Problem

  • Problem / challenge: Current methods for auditing text-to-image (T2I) generative AI models are limited in their ability to systematically explore the vast semantic space of AI-generated images. Auditors often rely on intuition, and existing tools lack structured methodologies for organizing and analyzing results.
  • Significance: Addressing biases, offensive content, and other risks in T2I models is critical for ensuring responsible AI deployment and mitigating societal harm.
  • Motivation and related work: Prior research has explored auditing tools for generative AI, focusing on text-based models or small-scale image evaluations. However, these tools struggle with scalability and fail to provide comprehensive support for navigating the semantic richness of images. There is a need for mixed-initiative tools that integrate human and AI capabilities to enable systematic and scalable auditing.

Solution

  • Proposed approach: Vipera, an interactive system combining visual guidance (scene graphs) and LLM-powered suggestions to support systematic auditing of T2I models.
  • Novelty:
    1. Integration of scene graphs to visually organize and summarize image semantics, enabling structured exploration.
    2. Use of LLM-powered suggestions to inspire new auditing criteria and prompts, addressing "unknown unknowns."
    3. Coordination of visual and LLM-driven guidance to reduce cognitive load and enhance systematic auditing.
    4. Tools for documenting and synthesizing findings into actionable audit reports.
  • Procedure and key techniques:
    • Scene graphs aggregate image semantics into a hierarchical structure with visual summaries (e.g., bar charts).
    • LLMs suggest new prompts and criteria based on differences between images and user-defined keywords.
    • Users interact with visual and textual guidance to iteratively refine their audits.
    • A note view integrates findings, bookmarks, and LLM-assisted auto-completion for report generation.

Results

  • Concrete findings:
    • Vipera reduced mental workload and improved performance compared to baseline systems, with significant gains in user-reported performance (p=0.0066).
    • Participants created more auditing criteria (87.2% and 78.4% of criteria in AI-supported systems) and explored more prompts and images.
    • AI guidance inspired broader exploration, while visual guidance encouraged systematic and detailed analysis.
  • Advantage over baselines:
    • Compared to systems without visual or AI support, Vipera enabled more systematic and thorough auditing with reduced cognitive load.
    • The combination of scene graphs and LLM suggestions outperformed either modality alone.
  • Experiments / evaluation:
    • Controlled study with 24 participants (20 general auditors, 4 experts) using four systems (baseline and three ablated versions of Vipera).
    • Metrics included NASA-TLX workload ratings, number of prompts, criteria, and bookmarks, and thematic analysis of audit reports.
  • Limitations and future work:
    • Study conducted in a controlled lab setting with student auditors; real-world and longitudinal evaluations are needed.
    • Fixed system order may introduce learning effects.
    • LLM inaccuracies occasionally misled participants; future work should improve label accuracy and user trust.
    • Opportunities for personalization, collaborative auditing, and fine-grained visual guidance remain unexplored.

Summary

Vipera is a novel system that combines visual scene graphs and LLM-powered suggestions to support systematic auditing of text-to-image generative AI models. Through a controlled user study, Vipera demonstrated its ability to reduce cognitive load, inspire broader exploration, and enable more systematic and thorough audits compared to baseline systems. The integration of visual and LLM-driven guidance proved complementary, enhancing both performance and user experience. Future work will focus on real-world deployments, improving AI label accuracy, and supporting collaborative auditing workflows.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222266/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791942
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Explainable AI (XAI), AI-Assisted Decision-Making & Automation, Algorithmic Transparency & Auditability, Interactive Data Visualization
work
Professions
AI/ML Researchers & Engineers, Data Scientists & Analysts, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers