Sensemaking in Multi-Agent LLM Interfaces: How Users Interpret Transparency and Trustworthiness Cues
Authors
Paper Title
Sensemaking in Multi-Agent LLM Interfaces: How Users Interpret Transparency and Trustworthiness Cues
Publication Info
- Topic area: User interpretation of transparency and trustworthiness in multi-agent Large Language Model (LLM) interfaces.
- Keywords: Multi-agent systems, transparency, trustworthiness, LLM interfaces, user mental models, epistemic cues, trust calibration, human-AI interaction, design-led study, interface design.
Background and Problem
- Problem / challenge: Current transparency mechanisms in AI systems often fail to align with how users interpret and engage with outputs of generative LLMs, especially in multi-agent settings. There is limited understanding of how users perceive and evaluate transparency and trustworthiness in such systems.
- Significance: Understanding user needs for transparency in multi-agent systems is crucial for designing interfaces that foster appropriate trust and support effective decision-making.
- Motivation and related work: Previous research has explored trust calibration in classical AI systems and initial studies on multi-agent LLMs, but these efforts remain scattered and lack a principled understanding of how users interpret multi-agent reasoning and transparency cues. This paper addresses this gap by focusing on user mental models and preferences for multi-agent transparency.
Solution
- Proposed approach: A design-led, qualitative, comparative structured observation study using five interface variants of multi-agent LLMs to explore user perceptions of transparency and trustworthiness.
- Novelty:
- A systematic exploration of how users interpret multi-agent reasoning and transparency cues.
- A reconceptualization of transparency as a context-sensitive sufficiency judgment rather than a volume dial.
- Identification of design tensions between visibility, interpretability, and cognitive effort.
- Introduction of progressive, on-demand transparency as a design strategy.
- Procedure and key techniques:
- Literature review to define a design space for multi-agent transparency, identifying seven design dimensions.
- Development of five interface variants (V1–V5) operationalizing different transparency configurations.
- In-person lab study with 12 participants, using think-aloud protocols, card-sorting activities, and semi-structured interviews.
- Analysis of user interactions with interfaces across two task types: information-seeking and logical reasoning.
Results
- Concrete findings:
- Users interpreted epistemic signals (e.g., disagreement, critique, consensus) as key cues for trustworthiness.
- Participants preferred a "Goldilocks" level of transparency, balancing informational value and cognitive effort.
- Task complexity and user expertise influenced transparency preferences, with simpler tasks requiring less visibility.
- Progressive, on-demand transparency was widely desired to manage cognitive workload.
- Advantage over baselines:
- Interfaces with agent-level rationales (e.g., V3) were perceived as more trustworthy and helpful compared to opaque designs (e.g., V1).
- Explicit critique (V4) and debate (V5) enhanced trust by showing the system "checking itself."
- Experiments / evaluation:
- Participants interacted with five interface variants across two tasks (information-seeking and reasoning).
- Evaluations included perceived transparency, helpfulness, and reliability using card-sorting activities.
- Data were analyzed using thematic analysis to identify user mental models and transparency preferences.
- Limitations and future work:
- Study tasks were predefined and bounded, limiting exploration of self-directed use cases.
- Interface configurations tested represent only a subset of the broader design space.
- Future work should examine transparency preferences in scenarios with agent errors and test progressive disclosure mechanisms at scale.
Summary
This study explores how users interpret transparency and trustworthiness in multi-agent LLM interfaces. By analyzing user interactions with five interface variants, the authors identify key epistemic cues (e.g., disagreement, critique, consensus) that shape trust perceptions. Transparency is reconceptualized as a context-sensitive sufficiency judgment, with users preferring progressive, on-demand transparency to balance cognitive effort and informational value. Task complexity, user expertise, and dispositional trust influence transparency needs. These findings provide actionable insights for designing trustworthy, human-centered multi-agent AI systems and highlight the need for future work on adaptive transparency mechanisms.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
Characterizing User-Reported Risks across LLM Chatbots
CHI '26· Human-LLM Collaboration +3
- 86%
AI and My Values: User Perceptions of LLMs’ Ability to Extract, Embody, and Explain Human Values from Casual Conversations
CHI '26· Human-LLM Collaboration +3
- 83%
Characterizing Unintended Consequences of GUI Agents For Web Browsing
CHI '26· Human-LLM Collaboration +2
- 71%
Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI Challenges
CHI '23· Human-LLM Collaboration +2
- 71%
Mind The Gap: Designers and Standards on Algorithmic System Transparency for Users
CHI '24· Explainable AI (XAI) +2
- 71%
Effects of LLM-based Search on Decision Making: Speed, Accuracy, and Overreliance
CHI '25· Human-LLM Collaboration +2
- 71%
Be Friendly, Not Friends: How LLM Sycophancy Shapes User Trust
CHI '26· Human-LLM Collaboration +2
- 71%
Personal Validation Effect in LLMs: Positive AI Responses Bias Perceptions of Validity, Reliability, Personalization, and Usefulness of Fictitious Predictions
CHI '26· Human-LLM Collaboration +2
- 71%
The AI Memory Gap: Users Misremember What They Created With AI or Without
CHI '26· Human-LLM Collaboration +2
- 71%
Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to Them
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)