Lexara: A User-Centered Toolkit for Evaluating Large Language Models for Conversational Visual Analytics

Human-LLM CollaborationInteractive Data VisualizationExplainable AI (XAI)Data Scientists & AnalystsAI/ML Researchers & EngineersSoftware Engineers & DevelopersHCI Researchers

Paper Title

Lexara: A User-Centered Toolkit for Evaluating Large Language Models for Conversational Visual Analytics

Publication Info

  • Topic area: Evaluation of Large Language Models (LLMs) for Conversational Visual Analytics (CVA).
  • Keywords: Conversational Visual Analytics, Large Language Models, evaluation toolkit, visualization quality, natural language metrics, multi-turn interactions, ambiguity resolution, user-centered design.

Background and Problem

  • Problem / challenge: Existing CVA evaluation methods are fragmented, require programming expertise, focus on single-turn interactions, and lack interpretable metrics for multi-format outputs (visualizations and text).
  • Significance: Evaluating LLMs for CVA is critical for ensuring trust, usability, and analytical accuracy in systems that democratize data exploration.
  • Motivation and related work: Prior CVA tools and benchmarks fail to capture real-world multi-turn, multi-format interactions and often overlook ambiguity and iterative refinement. Traditional NLP metrics struggle with CVA-specific outputs, and visualization-specific metrics lack comprehensive pipeline coverage.

Solution

  • Proposed approach: Lexara, a user-centered evaluation toolkit for CVA, operationalizes real-world insights into test cases, interpretable metrics, and an interactive low-code benchmarking tool.
  • Novelty:
    1. Real-world CVA test cases capturing multi-turn, multi-format interactions.
    2. Graded, interpretable metrics for visualization and natural language quality, accommodating multiple plausible answers.
    3. An interactive, low-code evaluation tool enabling systematic comparison of models and prompts.
  • Procedure and key techniques:
    • Semi-structured interviews with 22 CVA developers and observational studies with 16 end-users informed test case design.
    • Metrics evaluate visualization quality (data fidelity, semantic alignment, functional correctness, design clarity) and language quality (factual grounding, analytical reasoning, conversational coherence).
    • Interactive interface supports uploading datasources, configuring experiments, and inspecting multi-format outputs side-by-side.

Results

  • Concrete findings:
    • Lexara’s metrics align with human judgments (Spearman’s ρ = 0.68–0.82 for visualization metrics; ρ = 0.57–0.82 for natural language metrics).
    • Diary study participants conducted 38 experiments across 57 test cases, comparing 10 LLMs and 6 system prompts.
  • Advantage over baselines:
    • Lexara’s test cases and metrics capture real-world complexity better than synthetic benchmarks.
    • Graded correctness and multi-format evaluation provide nuanced insights absent in traditional NLP or visualization metrics.
  • Experiments / evaluation:
    • Validation study with human raters showed high inter-rater reliability (Cohen’s κ = 0.45–0.80).
    • Diary study participants valued Lexara’s interpretability, scalability, and support for multi-format comparisons.
  • Limitations and future work:
    • Current test suite coverage is limited to specific domains and common chart types.
    • YAML-based authoring poses challenges for non-technical users; future iterations may include point-and-click interfaces.
    • Metrics do not yet assess multimodal perception or tool-use capabilities of LLMs.

Summary

Lexara addresses critical gaps in evaluating LLMs for Conversational Visual Analytics by introducing real-world test cases, interpretable graded metrics, and an interactive low-code benchmarking tool. Validation studies demonstrate alignment between Lexara’s metrics and human judgments, while diary studies highlight its usability and diagnostic capabilities for practitioners. By enabling nuanced, scalable evaluation of multi-turn, multi-format CVA interactions, Lexara contributes to more trustworthy and user-centered development of LLM-based analytics systems. The toolkit is publicly available for broader adoption and extension.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/221866/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790735
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Human-LLM Collaboration, Interactive Data Visualization, Explainable AI (XAI)
work
Professions
Data Scientists & Analysts, AI/ML Researchers & Engineers, Software Engineers & Developers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers