Improving User Interface Generation Models from Designer Feedback

Human-LLM CollaborationPrototyping & User Testing360° Video & Panoramic ContentUI/UX DesignersAI/ML Researchers & EngineersHCI Researchers

Paper Title

Improving User Interface Generation Models from Designer Feedback

Publication Info

  • Topic area: Enhancing UI generation models using designer feedback and machine learning.
  • Keywords: UI generation, designer feedback, RLHF, sketching, commenting, revision, preference pairs, LLM fine-tuning, reward models, human-computer interaction.

Background and Problem

  • Problem / challenge: Current large language models (LLMs) struggle to generate well-designed UIs due to the absence of tacit design knowledge in training datasets. Existing reinforcement learning from human feedback (RLHF) methods, such as ranking and rating, are noisy and misaligned with designers' workflows.
  • Significance: Improving UI generation models can enable more efficient and high-quality UI design processes, benefiting designers and end-users.
  • Motivation and related work: Prior work has explored rubric-guided ratings, rankings, and annotation interfaces for collecting feedback, but these approaches often result in low inter-rater reliability and fail to capture nuanced design expertise. This paper builds on insights from designer workflows (e.g., commenting, sketching, revising) to propose a more effective feedback collection and model training methodology.

Solution

  • Proposed approach: A designer-aligned feedback collection and model training pipeline that incorporates commenting, sketching, and revising workflows to generate high-quality preference data for UI generation models.
  • Novelty:
    1. Techniques to transform designer comments, sketches, and revisions into machine-learnable preference pairs.
    2. A dataset of 1,460 UI screens annotated with designer feedback, showing reduced disagreement compared to traditional ranking methods.
    3. Validation of the approach through fine-tuning models, demonstrating improvements over baselines, including GPT-5.
  • Procedure and key techniques:
    • Generate synthetic UI descriptions and corresponding code using a base LLM (Qwen2.5-Coder 32B).
    • Collect designer feedback via four interfaces: ranking, commenting, sketching, and revising.
    • Convert feedback into preference pairs using automated pipelines.
    • Train reward models using a margin-based contrastive loss and fine-tune generator models using ORPO optimization.

Results

  • Concrete findings:
    • Models fine-tuned with sketch and revision feedback outperformed those trained with ranking or commenting data.
    • The best-performing model (Qwen3-Coder + Sketch) surpassed GPT-5 in human-judged UI quality.
    • Designer feedback led to a 61.7% agreement rate with HCI experts, with revision-based feedback achieving the highest agreement (76.1%).
  • Advantage over baselines:
    • Sketch-trained models achieved the highest performance, balancing data quality and quantity.
    • Fine-tuning with designer feedback enabled smaller models to outperform larger proprietary models like GPT-5.
  • Experiments / evaluation:
    • Designer feedback study with 21 participants generated 1,460 annotations.
    • Arena-style evaluations with six HCI experts compared models using Elo ratings and win rates.
    • Generalization evaluation showed consistent improvements across multiple base models.
  • Limitations and future work:
    • Limited validation with short-duration designer studies and synthetic UIs.
    • Need for broader evaluation with professional designs and larger participant pools.
    • Opportunities to explore other feedback types (e.g., usability studies) and adapt models for interactive UX evaluation.

Summary

This paper introduces a designer-aligned approach to improving UI generation models by leveraging feedback workflows like commenting, sketching, and revising. The resulting dataset and fine-tuned models demonstrate significant improvements over traditional ranking-based methods and outperform state-of-the-art proprietary models like GPT-5. The findings highlight the importance of high-quality, domain-specific feedback in training models and suggest future directions for integrating broader design expertise into machine learning workflows.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223499/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791567
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, Prototyping & User Testing, 360° Video & Panoramic Content
work
Professions
UI/UX Designers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers