How Do People Rank Multiple Mutant Agents?
Authors
Title of the Paper
How Do People Rank Multiple Mutant Agents?
Bibliographic Information
- Field of Study: Explainable Artificial Intelligence (XAI), Human-Computer Interaction
- Keywords: Explainable AI, After-Action Review, Ranking Task, Explanation Resolution, Mutant Agent Generation, Sequential Decision-Making, AI Evaluation
Research Background and Problem
- Identified Problem or Challenge: When faced with multiple AI-based decision systems simultaneously, how do people choose the most suitable system? More importantly, when a single system cannot fully meet the requirements, how can multiple systems be ranked as alternatives? Current explanation tools (XAI) face limitations in addressing complex sequential decision-making problems, especially when ranking multiple agents is required.
- Significance: Understanding these issues not only helps users better select and deploy AI systems but also has profound implications for advancing research in explainable AI. This is closely tied to improving user trust in AI and enhancing system transparency during use.
- Research Motivation and Related Work: The motivation is to provide a method that effectively supports users in ranking tasks, which is more operational and evaluative than existing XAI methods. This study references various related works, including cognitive models of explainable AI, decision explanations in gaming environments, and improvements in AI testing methods.
Solution
- Proposed Method or Solution:
- Explanation Resolution: Defines a quantitative metric to measure the effectiveness and granularity of explanations, enabling the identification of subtle performance differences between agents.
- Ranking Task: Proposes a new XAI experiential task that uses ranking tasks to evaluate explanation capabilities in multi-agent environments.
- Mutant Agent Generation: A technique based on multi-level noise perturbation of neural network weights to generate quality-controlled mutant agents.
- Innovations:
- Introduced a new metric for explainable AI called "Explanation Resolution," marking the first application of the resolution concept from microscopy to XAI.
- Designed a user-friendly task for agent ranking, integrating innovative explanation methods.
- Proposed a flexible and computationally low-cost mutant agent generation technique applicable to nearly all neural network-based systems.
- Implementation Steps and Key Techniques:
- Use the MNK game environment (an extended version of Tic-Tac-Toe) as the experimental domain.
- Train a base AI agent and generate a set of mutant agents.
- Design three types of explanation tools to showcase agent decision-making behavior:
- Scores Through-Time
- Scores On-the-Board
- Scores Best-to-Worst
- Conduct qualitative studies by analyzing user behavior and explanation choices during task completion.
Research Outcomes
- Specific Outcomes:
- Developed an integrated framework for explanation and task evaluation, including explanation tools and ranking tasks.
- Experimental findings revealed that users generally completed ranking tasks successfully, but higher explanation resolution was needed to distinguish between agents with subtle differences.
- Users exhibited diverse preferences for explanation methods, indicating that a single explanation type is insufficient for all users.
- Advantages Over Existing Solutions:
- The new method supports more flexible and fine-grained agent evaluation, especially when multiple agents have closely matched performance.
- The mutant agent generation technique is universal and low-cost, applicable to various AI models.
- Experimental or Evaluation Results:
- Ranking task results showed that users excelled at identifying top-ranked and bottom-ranked agents but made more errors with agents of similar mid-level performance.
- Users demonstrated a preference for diverse combinations of explanations, highlighting the importance of multiple explanation interactions.
- Limitations and Future Directions:
- The study was conducted solely in the MNK game environment, limiting its generalizability to other complex domains.
- Lack of comparative experiments prevents direct evaluation of whether the proposed explanation framework outperforms existing methods.
- Future work could explore the ecological validity of the mutation technique and its extension to real-world complex AI systems.
Summary and Implications
This study introduces a new method for quantifying explanation quality, "Explanation Resolution," combined with ranking tasks to enhance the effectiveness of XAI evaluation tasks. Additionally, the developed mutant agent generation technique provides a low-cost and flexible tool for systematically studying and improving XAI methods. Future research could validate these methods across more domains and explore optimal strategies for users when selecting combinations of explanation tools.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- When users need to evaluate multiple AI systems simultaneously, how can their explanation quality be quantified and compared?Category: Multi-System Explanation Comparison and EvaluationSimilar questionsarrow_forward
- In multi-agent settings, how can users effectively complete explanation-based ranking tasks?Category: Multi-System Explanation Comparison and EvaluationSimilar questionsarrow_forward
- What methods can generate quality-controlled AI variant agents to support user evaluation and research?Category: Multi-System Explanation Comparison and EvaluationSimilar questionsarrow_forward
Practical Problems
1- Users struggle to identify the most suitable among multiple AI systems and understand their decision processes.Category: Multi-System Explanation Comparison and EvaluationSimilar questionsarrow_forward
- 83%
What Did My Car Say? Impact of Autonomous Vehicle Explanation Errors and Driving Context On Comfort, Reliance, Satisfaction, and Driving Confidence
CHI '25· Automated Driving Interface & Takeover Design +2
- 83%
Mental Models in Human-AI Interaction: Systematic Review of Empirical Methodologies and Guidelines
IUI '26· Explainable AI (XAI) +2
- 80%
Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance
CHI '21· Explainable AI (XAI) +1
- 80%
One AI Does Not Fit All: A Cluster Analysis of the Laypeople’s Perception of AI Roles
CHI '23· Explainable AI (XAI) +1
- 80%
"Help Me Help the AI": Understanding How Explainability Can Support Human-AI Interaction
CHI '23· Explainable AI (XAI) +1
- 80%
Editable XAI: Toward Bidirectional Human-AI Alignment with Co-Editable Explanations of Interpretable Attributes
CHI '26· Explainable AI (XAI) +1
- 80%
Emergent, not Immanent: A Baradian Reading of Explainable AI
CHI '26· Explainable AI (XAI) +1
- 80%
Automated Rationale Generation: a Technique for Explainable AI and its Effects on Human Perceptions
IUI '19· Explainable AI (XAI) +1
- 80%
The Effects of Example-Based Explanations in a Machine Learning Interface
IUI '19· Explainable AI (XAI) +1
- 80%
Are Explanations Helpful? A Comparative Study of the Effects of Explanations in AI-Assisted Decision-Making
IUI '21· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)