Criticmate: Stagewise Human-AI Co-Critique in UI Design through Situation Awareness

Human-LLM CollaborationPrototyping & User TestingMobile App User ExperienceUI/UX DesignersAI/ML Researchers & EngineersHCI Researchers

Paper Title

Criticmate: Stagewise Human–AI Co-Critique in Single-Screen UI Evaluation

Publication Info

  • Topic area: Human–AI collaboration in UI evaluation
  • Keywords: UI evaluation, heuristic critique, human–AI collaboration, Situation Awareness, stagewise reasoning, single-screen evaluation, AI transparency, design critique, usability inspection, co-critique

Background and Problem

  • Problem / challenge: Existing AI tools for UI evaluation treat the process as a single-pass, black-box operation, limiting both AI reasoning and human involvement. This leads to generic, weakly grounded feedback and makes it difficult for humans to intervene effectively.
  • Significance: Improving the collaboration between humans and AI in UI evaluation can enhance critique quality, reduce cognitive load, and support better design decisions early in the lifecycle.
  • Motivation and related work: Prior work has explored heuristic evaluation, usability inspection, and AI-based tools for UI critique, but these approaches often lack structured reasoning and fail to integrate human expertise effectively. This paper builds on Situation Awareness (SA) theory to address these gaps.

Solution

  • Proposed approach: Criticmate, a stagewise human–AI co-critique system for single-screen UI evaluation, decomposes the evaluation process into three stages: Perception, Comprehension, and Projection, exposing intermediate reasoning artifacts for human intervention.
  • Novelty:
    1. Conceptualization of UI evaluation as a stagewise process grounded in SA theory.
    2. Implementation of Criticmate, an interactive system that operationalizes stagewise co-critique.
    3. Empirical validation of stagewise co-critique through offline benchmarks and a user study.
    4. Design implications for future human–AI evaluation tools.
  • Procedure and key techniques:
    1. Perception: AI parses the UI into sections and components, which humans validate and refine.
    2. Comprehension: AI describes visual and functional characteristics, with human adjustments for context and domain knowledge.
    3. Projection: AI and humans collaboratively identify gaps, propose fixes, and adapt guidelines to the specific task and domain.

Results

  • Concrete findings:
    • Criticmate’s stagewise pipeline achieved higher semantic similarity to expert critiques (F1@0.5 = 0.646) compared to single-pass baselines (e.g., ZS: 0.465, FS: 0.508).
    • It produced a more balanced mix of global and local critiques (≈39% local comments vs. <28% for baselines).
    • Generated more diverse critiques (Vendi Score: 3.91 for overall critiques, 6.07 for identified gaps).
  • Advantage over baselines:
    • Criticmate outperformed single-pass baselines in critique quality, engagement, and trust, as rated by both participants and external experts.
    • Enabled more concrete, consistent, and actionable feedback by exposing intermediate reasoning stages.
  • Experiments / evaluation:
    • Offline evaluation using the UICrit dataset (1,000 mobile UI screenshots) compared Criticmate to single-pass baselines and ablation variants.
    • Controlled user study with 26 participants (13 experts, 13 intermediates) compared Criticmate to a single-pass baseline in realistic critique workflows.
    • Metrics included semantic similarity, global–local balance, content diversity, and user engagement.
  • Limitations and future work:
    • Limited to single-screen, heuristic-style evaluation; does not address multi-screen flows or longitudinal critique.
    • Errors in earlier stages can propagate; future work could improve recognition accuracy and integrate visual annotations.
    • Study focused on professional designers and intermediates; further research needed with novices and non-design stakeholders.

Summary

Criticmate introduces a novel stagewise human–AI co-critique paradigm for single-screen UI evaluation, grounded in Situation Awareness theory. By structuring the process into Perception, Comprehension, and Projection stages, it enables transparent, editable reasoning and supports meaningful human intervention. Offline benchmarks and a user study demonstrated that Criticmate produces critiques that are more expert-like, balanced, and diverse than single-pass baselines, while fostering higher trust and engagement among users. This work highlights the potential of process-centric human–AI collaboration to improve both the quality of automated evaluation and the user experience of critique tools, with future opportunities to extend the approach to multi-screen and cross-device contexts.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222901/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790929
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, Prototyping & User Testing, Mobile App User Experience
work
Professions
UI/UX Designers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers