DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationExplainable AI (XAI)Computational Methods in HCISoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI Researchers

Paper Title

DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors

Publication Info

  • Topic area: Diagnosis and debugging of LLM-based multi-agent systems.
  • Keywords: LLM, multi-agent systems, debugging, diagnosis, Activity Theory, hierarchical summary, visualization, AI development, failure analysis, interactive tools.

Background and Problem

  • Problem / challenge: Diagnosing failures in LLM-based multi-agent systems is challenging due to their complexity, unstructured execution logs, and the diversity of agent behaviors. Existing tools often fail to provide high-level summaries or actionable insights for debugging.
  • Significance: Effective diagnosis is critical for improving multi-agent systems, which are increasingly used in complex domains such as web search, scientific research, and software development.
  • Motivation and related work: Current diagnostic tools (e.g., AutoGen Studio, LangSmith) focus on execution histories or visualizations but lack structured, multi-layered summaries that help developers identify root causes of failures. This paper addresses the gap by introducing a layered framework inspired by Activity Theory.

Solution

  • Proposed approach: DiLLS (Diagnosis of LLM-based Layered Summaries), an interactive system that organizes multi-agent behaviors into three hierarchical layers: activity, action, and operation.
  • Novelty:
    1. A three-layered framework based on Activity Theory to structure agent behaviors for diagnosis.
    2. Integration of natural language probing to extract hidden agent reasoning and progress.
    3. Interactive visualization tools to navigate and analyze multi-agent behaviors at different levels.
    4. Demonstrated effectiveness through a user study with AI developers.
  • Procedure and key techniques:
    • Extract and summarize agent behaviors into activity (high-level plans), action (goal-directed processes), and operation (low-level tasks).
    • Use natural language probing to elicit agents’ reasoning and progress.
    • Provide interactive visualizations: Activity View (overview of plans), Action View (execution steps), and Operation View (detailed logs).
    • Enable filtering and navigation to connect related behaviors across layers.

Results

  • Concrete findings:
    • Participants using DiLLS identified 3.67 failures on average, compared to 2.25 with the baseline system (p < 0.05).
    • Higher inter-participant agreement in failure identification with DiLLS (58% agreement for all participants vs. 20% with baseline).
    • Improved user confidence (average confidence score: 4.50 vs. 3.42, p < 0.05).
    • Reduced cognitive load across NASA-TLX dimensions, including mental demand, effort, and frustration.
  • Advantage over baselines:
    • Enhanced ability to detect high-level failures like problematic planning and action skipping.
    • Better support for tracing exploration processes and refining hypotheses.
    • Clearer connections between plans, actions, and operations.
  • Experiments / evaluation:
    • User study with 12 participants (6 male, 6 female, aged 24.17 ± 2.66 years) from academia and industry.
    • Tasks involved diagnosing errors in four cases from the GAIA benchmark using DiLLS and a baseline system.
    • Metrics: number of failures identified, confidence ratings, NASA-TLX scores, and qualitative feedback.
  • Limitations and future work:
    • Study focused only on failure diagnosis, not subsequent debugging or system modification.
    • Diagnosis results were subjective, relying on participants’ evidence and explanations.
    • Future work includes long-term studies, broader challenges in multi-agent behaviors, and deeper integration of Activity Theory constructs.

Summary

This paper introduces DiLLS, an interactive system for diagnosing failures in LLM-based multi-agent systems by organizing agent behaviors into three hierarchical layers: activity, action, and operation. Inspired by Activity Theory, the system uses natural language probing and interactive visualizations to help developers identify, understand, and analyze failures efficiently. A user study demonstrated that DiLLS significantly improves diagnostic accuracy, confidence, and cognitive load compared to baseline tools. The framework offers a practical solution for debugging multi-agent systems and lays the groundwork for future research on scalable and intuitive diagnostic tools.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223278/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790815
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, Explainable AI (XAI), Computational Methods in HCI
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers