SimUser: Generating Usability Feedback by Simulating Various Users Interacting with Mobile Applications
Authors
Title of the Paper
SimUser: Generating Usability Feedback by Simulating Various Users Interacting with Mobile Applications
Paper Information
- Research Domain: Human-Computer Interaction (HCI), User Experience (UX) Research, AI-driven User Simulation
- Keywords: Usability feedback, user simulation, large language models, mobile applications, prototyping, human-computer interaction, user characteristics, user scenarios, chain-of-thought (CoT), usability evaluation
Research Background and Problem Statement
- Identified Problems or Challenges:
- The conflict between rapid prototype iteration and time-consuming user testing.
- Existing AI-based methods focus on assessing system feasibility but often overlook the impact of user characteristics and usage contexts on usability.
- Limitations of datasets and the complex dynamic nature of usability issues make modeling interaction feedback for specific user groups difficult.
- Significance:
- Usability testing is critical for optimizing design, yet existing tools struggle to quickly and accurately simulate real interactions across diverse user groups.
- ISO 9241-11 (Human-System Interaction Standards) explicitly states that usability should account for the influence of specific user groups and their contexts.
- Research Motivation and Related Work:
- Current methods, such as visual saliency detection and interaction input prediction, fail to systematically consider the effects of user characteristics or context-related interactions.
- Although large language models (LLMs, such as GPT-4 and LLaMA) show potential in inferring user contexts and behaviors, they still face limitations in understanding interfaces and user perceptions.
Solution
Methodology and Innovations
- Method: Develop an LLM-based tool, SimUser, combining Chain-of-Thought (CoT) reasoning and user modeling techniques to generate usability feedback by simulating user interactions with applications.
- Utilize two sub-agents (Mobile Application Agent and User Agent) to represent the mobile application and simulated user, respectively.
- Introduce detailed user interaction contexts through user characteristic modeling and scenario expansion.
- Employ the concept of "Expectation Disconfirmation" to identify usability issues.
- Innovations:
- Emphasizes the impact of user characteristics (e.g., cognitive ability, interaction skills) and contextual factors on usability feedback.
- Leverages CoT reasoning to reduce biases during the inference process through step-by-step reasoning and control.
- Proposes a detailed simulation execution process: the Mobile Application Agent handles interface descriptions and logical feedback, while the User Agent generates user expectations, interaction simulations, and feedback.
Implementation Steps
- Input Stage: Designers upload prototype code and interface images, and define target user characteristics and testing tasks.
- Mobile Application Agent:
- Generates natural language descriptions of the interface, including layout, contrast, visual saliency, etc.
- Defines operational logic: constructs a structural logic map of the interface based on code files.
- User Agent:
- Creates user profiles based on user characteristics and defines detailed attributes (e.g., visual needs, operational proficiency).
- Infers simulated user expectations of the interface, simulates user operations step-by-step, and generates interaction feedback.
- Interaction Process:
- The User Agent predicts interface content through human behavior patterns.
- Simulates user perceptions, errors, and subjective evaluations after task completion.
- Feedback Output:
- Provides usability feedback for individual interfaces and overall interaction logic.
Research Results
- Specific Findings:
- SimUser achieved coverage rates ranging from 35.7% to 100% in simple smartwatch application tests, with an average coverage rate of 80% for human usability feedback.
- Generated richer contextual scenarios and feedback compared to human users (e.g., extreme scenarios and edge cases).
- Covered over 70% of user scenarios during simulated interactions.
- Experimental results demonstrated high consistency with real human users in terms of interface information, interaction logic, and operational feasibility.
- Advantages:
- Compared to traditional user testing, SimUser is fast, low-cost, and capable of simulating multiple complex scenarios.
- Provides abundant design insights, aiding usability analysis during prototype iteration.
- Experiments or Evaluation:
- Experiments compared feedback from real users (college students and elderly users) with SimUser-generated feedback.
- Collected evaluations from 24 human users and 21 designers regarding SimUser. The average SUS (System Usability Scale) score was 66, close to the standard average.
- Limitations and Future Directions:
- Limitations:
- Compatibility testing for complex interfaces (e.g., smartphones) is not yet fully covered.
- SimUser relies on complete user profiles and environmental information, which may lead to high input material requirements.
- Excessive redundant feedback data (not mentioned by human users) increases the burden on designers to filter information.
- Future Directions:
- Enhance LLM capabilities to simulate emotions and dynamic factors, making simulation results more aligned with subtle user experience differences.
- Expand to more complex design scenarios, such as multimodal interaction interfaces (e.g., in-vehicle interfaces).
- Explore integration of generated feedback directly into design tools to improve usability and interpretability.
- Limitations:
The above summary comprehensively outlines the core findings and academic contributions of this research, while providing directions for future studies and practical product optimization.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can usability feedback be generated by simulating users with different characteristics interacting with mobile apps?Category: AI/LLM as Design Collaborators and Creative ToolsSimilar questionsarrow_forward
- How do user characteristics (e.g., cognitive ability and interaction skill) and contextual factors affect usability feedback generation?Category: AI/LLM as Design Collaborators and Creative ToolsSimilar questionsarrow_forward
- Can LLMs with chain-of-thought (CoT) techniques improve the coverage and precision of usability feedback?Category: AI/LLM as Design Collaborators and Creative ToolsSimilar questionsarrow_forward
Practical Problems
1- Designers struggle to quickly obtain accurate usability feedback across multiple users and scenarios.Category: AI/LLM as Design Collaborators and Creative ToolsSimilar questionsarrow_forward
- 75%
Generating Automatic Feedback on UI Mockups with Large Language Models
CHI '24· Human-LLM Collaboration +1
- 75%
Enhancing UX Evaluation Through Collaboration with Conversational AI Assistants: Effects of Proactive Dialogue and Timing
CHI '24· Human-LLM Collaboration +2
- 75%
PromptInfuser: How Tightly Coupling AI and UI Design Impacts Designers’ Workflows
DIS '24· Human-LLM Collaboration +1
- 75%
What About My Design Context?: Exploring the Use of Generative AI to Support Customization of Translational Research Artifacts
DIS '25· Human-LLM Collaboration +1
- 75%
JAMPL8: Exploring LLM-Enhanced Templates for Idea Reflection
IUI '24· Human-LLM Collaboration +1
- 75%
StoryEnsemble: Enabling Dynamic Exploration & Iteration in the Design Process with AI and Forward-Backward Propagation
UIST '25· Human-LLM Collaboration +1
- 67%
Using High Frequency Accelerometer and Mouse to Compensate for End-to-end Latency in Indirect Interaction
CHI '18· Prototyping & User Testing
- 60%
May AI? Design Ideation with Cooperative Contextual Bandits
CHI '19· Generative AI (Text, Image, Music, Video) +2
- 60%
'It Is Not Always Discovery Time': Four Pragmatic Approaches in Designing AI Systems
CHI '22· Human-LLM Collaboration +2
- 60%
Jigsaw: Supporting Designers to Prototype Multimodal Applications by Chaining AI Foundation Models
CHI '24· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)