Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI
Authors
Paper Title
Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI
Publication Info
- Topic area: Systematic auditing of text-to-image generative AI models
- Keywords: Text-to-image models, AI auditing, scene graph, LLM guidance, systematic auditing, visual analytics, generative AI, human-AI collaboration, bias detection, responsible AI
Background and Problem
- Problem / challenge: Current methods for auditing text-to-image (T2I) generative AI models are limited in their ability to systematically explore the vast semantic space of AI-generated images. Auditors often rely on intuition, and existing tools lack structured methodologies for organizing and analyzing results.
- Significance: Addressing biases, offensive content, and other risks in T2I models is critical for ensuring responsible AI deployment and mitigating societal harm.
- Motivation and related work: Prior research has explored auditing tools for generative AI, focusing on text-based models or small-scale image evaluations. However, these tools struggle with scalability and fail to provide comprehensive support for navigating the semantic richness of images. There is a need for mixed-initiative tools that integrate human and AI capabilities to enable systematic and scalable auditing.
Solution
- Proposed approach: Vipera, an interactive system combining visual guidance (scene graphs) and LLM-powered suggestions to support systematic auditing of T2I models.
- Novelty:
- Integration of scene graphs to visually organize and summarize image semantics, enabling structured exploration.
- Use of LLM-powered suggestions to inspire new auditing criteria and prompts, addressing "unknown unknowns."
- Coordination of visual and LLM-driven guidance to reduce cognitive load and enhance systematic auditing.
- Tools for documenting and synthesizing findings into actionable audit reports.
- Procedure and key techniques:
- Scene graphs aggregate image semantics into a hierarchical structure with visual summaries (e.g., bar charts).
- LLMs suggest new prompts and criteria based on differences between images and user-defined keywords.
- Users interact with visual and textual guidance to iteratively refine their audits.
- A note view integrates findings, bookmarks, and LLM-assisted auto-completion for report generation.
Results
- Concrete findings:
- Vipera reduced mental workload and improved performance compared to baseline systems, with significant gains in user-reported performance (p=0.0066).
- Participants created more auditing criteria (87.2% and 78.4% of criteria in AI-supported systems) and explored more prompts and images.
- AI guidance inspired broader exploration, while visual guidance encouraged systematic and detailed analysis.
- Advantage over baselines:
- Compared to systems without visual or AI support, Vipera enabled more systematic and thorough auditing with reduced cognitive load.
- The combination of scene graphs and LLM suggestions outperformed either modality alone.
- Experiments / evaluation:
- Controlled study with 24 participants (20 general auditors, 4 experts) using four systems (baseline and three ablated versions of Vipera).
- Metrics included NASA-TLX workload ratings, number of prompts, criteria, and bookmarks, and thematic analysis of audit reports.
- Limitations and future work:
- Study conducted in a controlled lab setting with student auditors; real-world and longitudinal evaluations are needed.
- Fixed system order may introduce learning effects.
- LLM inaccuracies occasionally misled participants; future work should improve label accuracy and user trust.
- Opportunities for personalization, collaborative auditing, and fine-grained visual guidance remain unexplored.
Summary
Vipera is a novel system that combines visual scene graphs and LLM-powered suggestions to support systematic auditing of text-to-image generative AI models. Through a controlled user study, Vipera demonstrated its ability to reduce cognitive load, inspire broader exploration, and enable more systematic and thorough audits compared to baseline systems. The integration of visual and LLM-driven guidance proved complementary, enhancing both performance and user experience. Future work will focus on real-world deployments, improving AI label accuracy, and supporting collaborative auditing workflows.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
Gamut: A Design Probe to Understand How Data Scientists Understand Machine Learning Models
CHI '19· Explainable AI (XAI) +2
- 86%
User Ex Machina : Simulation as a Design Probe in Human-in-the-Loop Text Analytics
CHI '21· Explainable AI (XAI) +3
- 86%
RELIC: Investigating Large Language Model Responses using Self-Consistency
CHI '24· Explainable AI (XAI) +2
- 86%
Comparables XAI: Faithful Example-based AI Explanations with Counterfactual Trace Adjustments
CHI '26· Explainable AI (XAI) +2
- 86%
Transferable XAI: Relating Understanding Across Domains with Explanation Transfer
IUI '26· Explainable AI (XAI) +2
- 75%
Drava: Aligning Human Concepts with Machine Learning Latent Dimensions for the Visual Exploration of Small Multiples
CHI '23· Explainable AI (XAI) +2
- 75%
From Overload to Convergence: Supporting Multi-Issue Human–AI Negotiation with Bayesian Visualization
CHI '26· Explainable AI (XAI) +3
- 75%
From Reflection to Repair: A Scoping Review of Dataset Documentation Tools
CHI '26· Explainable AI (XAI) +3
- 75%
StepMIND: A Visual Framework for Stepwise, Multimodal, and Bidirectional Explanations of AI-Generated Data Analysis Pipeline
IUI '26· Explainable AI (XAI) +3
- 71%
Explanations as Mechanisms for Supporting Algorithmic Transparency
CHI '18· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)