Improving User Interface Generation Models from Designer Feedback
Authors
Paper Title
Improving User Interface Generation Models from Designer Feedback
Publication Info
- Topic area: Enhancing UI generation models using designer feedback and machine learning.
- Keywords: UI generation, designer feedback, RLHF, sketching, commenting, revision, preference pairs, LLM fine-tuning, reward models, human-computer interaction.
Background and Problem
- Problem / challenge: Current large language models (LLMs) struggle to generate well-designed UIs due to the absence of tacit design knowledge in training datasets. Existing reinforcement learning from human feedback (RLHF) methods, such as ranking and rating, are noisy and misaligned with designers' workflows.
- Significance: Improving UI generation models can enable more efficient and high-quality UI design processes, benefiting designers and end-users.
- Motivation and related work: Prior work has explored rubric-guided ratings, rankings, and annotation interfaces for collecting feedback, but these approaches often result in low inter-rater reliability and fail to capture nuanced design expertise. This paper builds on insights from designer workflows (e.g., commenting, sketching, revising) to propose a more effective feedback collection and model training methodology.
Solution
- Proposed approach: A designer-aligned feedback collection and model training pipeline that incorporates commenting, sketching, and revising workflows to generate high-quality preference data for UI generation models.
- Novelty:
- Techniques to transform designer comments, sketches, and revisions into machine-learnable preference pairs.
- A dataset of 1,460 UI screens annotated with designer feedback, showing reduced disagreement compared to traditional ranking methods.
- Validation of the approach through fine-tuning models, demonstrating improvements over baselines, including GPT-5.
- Procedure and key techniques:
- Generate synthetic UI descriptions and corresponding code using a base LLM (Qwen2.5-Coder 32B).
- Collect designer feedback via four interfaces: ranking, commenting, sketching, and revising.
- Convert feedback into preference pairs using automated pipelines.
- Train reward models using a margin-based contrastive loss and fine-tune generator models using ORPO optimization.
Results
- Concrete findings:
- Models fine-tuned with sketch and revision feedback outperformed those trained with ranking or commenting data.
- The best-performing model (Qwen3-Coder + Sketch) surpassed GPT-5 in human-judged UI quality.
- Designer feedback led to a 61.7% agreement rate with HCI experts, with revision-based feedback achieving the highest agreement (76.1%).
- Advantage over baselines:
- Sketch-trained models achieved the highest performance, balancing data quality and quantity.
- Fine-tuning with designer feedback enabled smaller models to outperform larger proprietary models like GPT-5.
- Experiments / evaluation:
- Designer feedback study with 21 participants generated 1,460 annotations.
- Arena-style evaluations with six HCI experts compared models using Elo ratings and win rates.
- Generalization evaluation showed consistent improvements across multiple base models.
- Limitations and future work:
- Limited validation with short-duration designer studies and synthetic UIs.
- Need for broader evaluation with professional designs and larger participant pools.
- Opportunities to explore other feedback types (e.g., usability studies) and adapt models for interactive UX evaluation.
Summary
This paper introduces a designer-aligned approach to improving UI generation models by leveraging feedback workflows like commenting, sketching, and revising. The resulting dataset and fine-tuned models demonstrate significant improvements over traditional ranking-based methods and outperform state-of-the-art proprietary models like GPT-5. The findings highlight the importance of high-quality, domain-specific feedback in training models and suggest future directions for integrating broader design expertise into machine learning workflows.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
A study of UX Practitioners Roles in Designing Real-World, Enterprise ML Systems
CHI '22· Human-LLM Collaboration +2
- 71%
Design Principles for Generative AI Applications
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 71%
Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-Oz
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
Automating UI Optimization through Multi-Agentic Reasoning
CHI '26· Human-LLM Collaboration +2
- 71%
Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces
CHI '26· Human-LLM Collaboration +2
- 71%
Criticmate: Stagewise Human-AI Co-Critique in UI Design through Situation Awareness
CHI '26· Human-LLM Collaboration +2
- 71%
When Designers Sweat: Behavioral Traces of GenAI Co-Creation
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
Orality: A Semantic Canvas for Externalizing and Clarifying Thoughts with Speech
CHI '26· Human-LLM Collaboration +2
- 67%
Mapping Machine Learning Advances from HCI Research to Reveal Starting Places for Design Innovation
CHI '18· Human-LLM Collaboration
Based on Jaccard similarity of research subtopics & professions (≥60%)