DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
Authors
Paper Title
DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
Publication Info
- Topic area: Diagnosis and debugging of LLM-based multi-agent systems.
- Keywords: LLM, multi-agent systems, debugging, diagnosis, Activity Theory, hierarchical summary, visualization, AI development, failure analysis, interactive tools.
Background and Problem
- Problem / challenge: Diagnosing failures in LLM-based multi-agent systems is challenging due to their complexity, unstructured execution logs, and the diversity of agent behaviors. Existing tools often fail to provide high-level summaries or actionable insights for debugging.
- Significance: Effective diagnosis is critical for improving multi-agent systems, which are increasingly used in complex domains such as web search, scientific research, and software development.
- Motivation and related work: Current diagnostic tools (e.g., AutoGen Studio, LangSmith) focus on execution histories or visualizations but lack structured, multi-layered summaries that help developers identify root causes of failures. This paper addresses the gap by introducing a layered framework inspired by Activity Theory.
Solution
- Proposed approach: DiLLS (Diagnosis of LLM-based Layered Summaries), an interactive system that organizes multi-agent behaviors into three hierarchical layers: activity, action, and operation.
- Novelty:
- A three-layered framework based on Activity Theory to structure agent behaviors for diagnosis.
- Integration of natural language probing to extract hidden agent reasoning and progress.
- Interactive visualization tools to navigate and analyze multi-agent behaviors at different levels.
- Demonstrated effectiveness through a user study with AI developers.
- Procedure and key techniques:
- Extract and summarize agent behaviors into activity (high-level plans), action (goal-directed processes), and operation (low-level tasks).
- Use natural language probing to elicit agents’ reasoning and progress.
- Provide interactive visualizations: Activity View (overview of plans), Action View (execution steps), and Operation View (detailed logs).
- Enable filtering and navigation to connect related behaviors across layers.
Results
- Concrete findings:
- Participants using DiLLS identified 3.67 failures on average, compared to 2.25 with the baseline system (p < 0.05).
- Higher inter-participant agreement in failure identification with DiLLS (58% agreement for all participants vs. 20% with baseline).
- Improved user confidence (average confidence score: 4.50 vs. 3.42, p < 0.05).
- Reduced cognitive load across NASA-TLX dimensions, including mental demand, effort, and frustration.
- Advantage over baselines:
- Enhanced ability to detect high-level failures like problematic planning and action skipping.
- Better support for tracing exploration processes and refining hypotheses.
- Clearer connections between plans, actions, and operations.
- Experiments / evaluation:
- User study with 12 participants (6 male, 6 female, aged 24.17 ± 2.66 years) from academia and industry.
- Tasks involved diagnosing errors in four cases from the GAIA benchmark using DiLLS and a baseline system.
- Metrics: number of failures identified, confidence ratings, NASA-TLX scores, and qualitative feedback.
- Limitations and future work:
- Study focused only on failure diagnosis, not subsequent debugging or system modification.
- Diagnosis results were subjective, relying on participants’ evidence and explanations.
- Future work includes long-term studies, broader challenges in multi-agent behaviors, and deeper integration of Activity Theory constructs.
Summary
This paper introduces DiLLS, an interactive system for diagnosing failures in LLM-based multi-agent systems by organizing agent behaviors into three hierarchical layers: activity, action, and operation. Inspired by Activity Theory, the system uses natural language probing and interactive visualizations to help developers identify, understand, and analyze failures efficiently. A user study demonstrated that DiLLS significantly improves diagnostic accuracy, confidence, and cognitive load compared to baseline tools. The framework offers a practical solution for debugging multi-agent systems and lays the groundwork for future research on scalable and intuitive diagnostic tools.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 88%
Towards Guidelines for Designing Human-in-the-Loop Machine Training Interfaces
IUI '21· Human-LLM Collaboration +3
- 86%
PaTAT: Human-AI Collaborative Qualitative Coding with Explainable Interactive Rule Synthesis
CHI '23· Explainable AI (XAI) +2
- 86%
Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs
CHI '23· Explainable AI (XAI) +2
- 86%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 86%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 86%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
- 86%
Unakite: Scaffolding Developers’ Decision-Making Using the Web
UIST '19· Explainable AI (XAI) +2
- 75%
When Help Hurts: Verification Load and Fatigue with AI Coding Assistants
CHI '26· Human-LLM Collaboration +3
- 71%
Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making
CHI '23· Human-LLM Collaboration +1
- 71%
Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts
CHI '23· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)