Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework
Authors
Paper Title
Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework
Publication Info
- Topic area: Chart data extraction using multimodal large language models (MLLMs).
- Keywords: Chart data extraction, multimodal large language models, ExChart, ExChart-Bench, coordinate system perception, chart-to-table alignment, numerical accuracy, mixed-initiative workflows, visualization, human-computer interaction.
Background and Problem
- Problem / challenge: Current chart data extraction methods, including fully automatic, interactive, and mixed-initiative systems, struggle with efficiency, generalizability, and numerical accuracy, particularly for charts without visible data labels. MLLMs offer potential but are unreliable for precise value recovery.
- Significance: Reliable chart data extraction is critical for reproducibility, analysis, and accessibility in fields like research, finance, and policy. Inaccurate extraction can lead to flawed analyses and decisions.
- Motivation and related work: Existing systems like ChartOCR and WebPlotDigitizer are either brittle or tedious. Mixed-initiative systems improve efficiency but lack generalizability. MLLMs provide a unified interface but fail to accurately infer values without labels. This paper addresses these gaps by focusing on improving MLLM performance for chart data extraction.
Solution
- Proposed approach: The authors introduce ExChart, a two-stage training framework inspired by human chart-reading processes, and ExChart-Bench, a benchmark for evaluating chart data extraction performance.
- Novelty:
- Development of ExChart-Bench, a benchmark featuring real-world and synthetic charts without data labels.
- Introduction of a two-stage training framework (Coordinate System Perception Enhancement and Chart-Table Alignment) to improve MLLM performance.
- Demonstration of state-of-the-art performance with a lightweight 7B-parameter model.
- Validation of the framework's effectiveness in mixed-initiative workflows through a user study.
- Procedure and key techniques:
- Coordinate System Perception Enhancement (CSPE): Trains models to understand coordinate geometry and visual encodings.
- Chart-Table Alignment (CTA): Fine-tunes models to generate structured data tables with high numerical fidelity.
- Construction of ExChart-Bench: Includes 3,600 chart-table-prompt pairs across diverse chart types and styles, with Adaptive MAPE as the primary evaluation metric.
- User study to evaluate the integration of MLLMs into interactive workflows.
Results
- Concrete findings:
- ExChart achieves an Adaptive MAPE of 4.87%, significantly outperforming the base model (16.01%) and larger models like GLM-4.5V (5.94%).
- Format success rate improves from 89.58% to 99.11%.
- Substantial improvements in handling polar coordinate charts (e.g., radar charts: Adaptive MAPE reduced from 25.47% to 3.21%).
- Advantage over baselines:
- ExChart outperforms state-of-the-art models like GLM-4.5V and Gemini 2.5 Flash in both numerical accuracy and structured output generation.
- Demonstrates superior performance compared to chart-specific models (e.g., ChartInstruct, UniChart) on RNSS and RMS metrics.
- Experiments / evaluation:
- Benchmarked 11 MLLMs across 3,600 charts using ExChart-Bench.
- Conducted ablation studies to validate the contributions of CSPE and CTA stages.
- User study with 12 participants showed that MLLM-assisted workflows enable efficient and reliable chart data extraction, with an average SUS score of 94.79.
- Limitations and future work:
- Current models, even with ExChart, are not yet reliable enough for fully automatic extraction.
- Future work includes expanding chart diversity, exploring table shape impacts, and incorporating uncertainty estimation for enhanced interpretability.
Summary
This paper addresses the limitations of MLLMs in chart data extraction by introducing ExChart, a human-inspired two-stage training framework, and ExChart-Bench, a benchmark for systematic evaluation. The proposed approach significantly improves numerical accuracy and structured output generation, achieving state-of-the-art performance with a lightweight 7B-parameter model. A user study confirms the effectiveness of integrating MLLMs into mixed-initiative workflows, enabling efficient and reliable chart data extraction. These findings highlight the potential of combining tailored training strategies with human-AI collaboration for practical applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Method for Exploring Generative Adversarial Networks (GANs) via Automatically Generated Image Galleries
CHI '21· Generative AI (Text, Image, Music, Video) +2
- 83%
Flowco: Mixed-Initiative Authoring of Reliable End-to-End Data Analyses via Dataflow Graphs and LLMs
UIST '25· Generative AI (Text, Image, Music, Video) +2
- 71%
Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs
CHI '25· Human-LLM Collaboration +2
- 71%
Dango: A Mixed-Initiative Data Wrangling System using Large Language Model
CHI '25· Human-LLM Collaboration +2
- 71%
PleaSQLarify: Visual Pragmatic Repair for Natural Language Database Querying
CHI '26· Human-LLM Collaboration +2
- 71%
SCSimulator: An Exploratory Visual Analytics Framework for Partner Selection in Supply Chains through LLM-driven Multi-Agent Simulation
IUI '26· Human-LLM Collaboration +2
- 71%
Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition
UIST '24· Human-LLM Collaboration +2
- 71%
Reality Proxy: Fluid Interactions with Real-World Objects in MR via Abstract Representations
UIST '25· Mixed Reality Workspaces +2
- 67%
Visualizing Examples of Deep Neural Networks at Scale
CHI '21· Human-LLM Collaboration +1
- 67%
Discovering the Syntax and Strategies of Natural Language Programming with Generative Language Models
CHI '22· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)