GANzilla: User-Driven Direction Discovery in Generative Adversarial Networks
Authors
Title of the Paper
GANzilla: User-Driven Direction Discovery in Generative Adversarial Networks
Paper Information
- Domain: Human-Computer Interaction, Generative Adversarial Networks (GAN), Explainable AI
- Keywords: Generative Adversarial Networks, Editing Directions, Interactive System, Explainable AI, Image Editing, User-Driven, User Experience, Deep Learning, Facial Editing, Style Direction Exploration
Research Background and Problem
- Problem or Challenge:
- Although GANs (Generative Adversarial Networks) have demonstrated remarkable performance in various fields (e.g., image stylization, scene generation), they remain a "black box" for end-users, lacking transparency and effective control over the characteristics of generated images.
- Existing methods are mostly algorithm-driven (e.g., principal component analysis or semantic control), making it difficult for users to actively define personalized editing directions.
- Importance:
- Providing transparency and control over the generation process is crucial for promoting the practical application of GANs, especially for non-expert users with specific creative needs.
- Expanding the application of GANs from "expert-only" to broader user groups such as designers and artists.
- Research Motivation and Related Work:
- Previous research has explored algorithm-driven direction extraction and predefined controls, but methods for user-driven discovery of editing directions remain insufficient.
- This study builds on traditional scatter/gather interaction techniques to develop GANzilla, a user-centered tool for direction discovery.
Solution
- Basic Approach:
- Introduce a tool named GANzilla that helps users discover GAN editing directions that meet their goals through interaction (e.g., scatter/gather techniques) without relying on predefined directions.
- Built on StyleGAN2, the workflow is designed to be generalizable to other GAN models.
- Innovations:
- User-driven instead of algorithm-driven: Users actively participate in the direction discovery process through filtering and iterative optimization.
- Integration of classic interaction techniques (e.g., user-friendly scatter/gather mechanisms) to support extensive exploration and selection.
- Implementation Steps and Key Techniques:
- Highlight Region Selection (optional): Users mark the desired image region for editing using a brush tool.
- Direction Sampling and Clustering: Based on user markings (if any), a large number of directions are randomly sampled from the StyleSpace of StyleGAN2 and presented as thumbnails via K-means clustering for user selection.
- Iterative Scatter/Gather: Users filter interesting directions, explore new directions through scatter plots, or return to previous clustering nodes for further selection.
- Direction Testing: Users test the effectiveness of editing directions on multiple images and adjust the intensity of directions using sliders.
- Direction Bookmarking: Users can save satisfactory directions for future use.
- Technical Implementation:
- Backend utilizes StyleGAN2 to generate directions.
- The complete deep learning and front-end user interface are built using frameworks such as PyTorch, Flask, and React.
Research Outcomes
- Specific Results:
- In closed tasks, GANzilla helps users generate directions that edit reference images into content similar to target images, outperforming random direction samples.
- In open-ended tasks, users utilized GANzilla to achieve personalized image edits (e.g., making faces appear happier, older, or surprised), with editing results showing minimal semantic distance from task objectives.
- Experimental or Evaluation Results:
- Closed Tasks: Images generated using user-discovered directions significantly outperformed those from random directions in terms of target similarity; in 33 out of 36 tasks, user-generated results ranked among the top 5 of random samples.
- Open-Ended Tasks:
- Using DeepFace, user-generated age/emotion editing directions showed an average age increase of about 10 years and emotion indicators improved by over 50%.
- Text-image semantic similarity analysis based on the CLIP model revealed that user-generated directions were closer to textual descriptions such as "old," "happy," and "surprised" compared to random directions.
- User Behavior: Users tended to test directions multiple times in open-ended tasks, while relying more on scatter/gather techniques in closed tasks, indicating that task type influences user behavior patterns.
- User Feedback:
- Overall usability was rated high, with users finding the tool's learning curve low.
- Key components (e.g., scatter/gather, multi-image direction testing) received high praise.
- Some users suggested improving GAN direction interaction (e.g., reducing randomness, enhancing recommendation features).
- Limitations and Future Directions:
- Research Limitations: Lack of large-scale user experiments; fixed task sequence (open-ended followed by closed); a few participants had prior GAN knowledge.
- Improvement Directions:
- Enhance GAN disentanglement (reduce unintended coupling of direction edits, e.g., "smile" causing "youthfulness").
- Provide more data visualization and user guidance (e.g., heatmap display of direction change paths).
- Combine algorithm-driven direction discovery methods to complement user-driven strategies.
Summary
This paper presents an intuitive user interaction platform (GANzilla), where the user-driven approach complements the typically algorithm-centric direction exploration models in GAN applications, bringing greater transparency and control to complex generative models.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do users discover editing directions in GAN models through the GANzilla interactive system?Category: Exploratory Image Retrieval and Generation ControlSimilar questionsarrow_forward
- How does user-driven direction discovery perform in generative image editing compared with algorithm-driven approaches?Category: Exploratory Image Retrieval and Generation ControlSimilar questionsarrow_forward
- How is user behavior affected by task type (open-ended or closed-ended)?Category: Exploratory Image Retrieval and Generation ControlSimilar questionsarrow_forward
Practical Problems
1- Lay users struggle to transparently and efficiently control characteristics of GAN-generated images.Category: Exploratory Image Retrieval and Generation ControlSimilar questionsarrow_forward
- 83%
AI-Augmented Brainwriting: Investigating the use of LLMs in group ideation
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 83%
Design Principles for Generative AI Applications
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 83%
Preference-Guided Prompt Optimization for Text-to-Image Generation
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 83%
Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-Oz
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 83%
Partnering with Generative AI: Experimental Evaluation of Model-Led and Human-Led Interaction in Human-AI Co-Creation
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 83%
Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 83%
When Designers Sweat: Behavioral Traces of GenAI Co-Creation
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 83%
Adaptive Prompt Elicitation for Text-to-Image Generation
IUI '26· Generative AI (Text, Image, Music, Video) +2
- 80%
Mapping Machine Learning Advances from HCI Research to Reveal Starting Places for Design Innovation
CHI '18· Human-LLM Collaboration
- 80%
OPTIMISM: Enabling Collaborative Implementation of Domain-Specific Metaheuristic Optimization
CHI '23· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)