Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs
Authors
Research Background and Problem Statement
-
Problems or Challenges:
- In Exploratory Data Analysis (EDA) and data storytelling, extracting and conveying actionable insights from complex data is an important yet challenging task.
- Actionable insights require not only data-driven facts but also the adjustment of analytical strategies based on real-world decision-making needs and the integration of domain knowledge.
- Existing tools fall short in generating actionable insights and delivering detailed and persuasive analytical results. For instance, while large language models (LLMs) can generate narratives, they may lack interpretability and contextual precision.
-
Significance:
- Actionable insights guide effective decision-making and concrete actions, which are core needs in fields such as business analytics, educational data mining, and public health.
- Leveraging large language models can further enhance the generation, articulation, and integration of analytical results with domain knowledge, facilitating more effective decision-making processes.
-
Research Motivation: The authors aim to optimize the EDA and data storytelling process by developing an LLM-powered tool that achieves semantic precision, rhetorical persuasiveness, and practical relevance, helping users overcome the primary challenges in data analysis and insight communication.
Solution
-
Proposed Method or Solution: The authors propose a design space that encompasses three key dimensions of EDA and data storytelling:
- Semantic Precision: Ensuring the accuracy and clarity of analytical results.
- Rhetorical Persuasion: Enhancing the persuasiveness of results through strategic analysis and narrative.
- Pragmatic Relevance: Integrating data-driven facts with domain knowledge to form actionable insights.
Based on this design space, the authors developed "Jupybara," an AI assistant implemented as a Jupyter Notebook extension. It employs the following two key strategies:
- Design-Space-Aware Prompting: Translating the theoretical framework into specific prompts to guide the LLM in optimizing each dimension.
- Multi-Agent Architecture: The system leverages multiple agents performing distinct tasks (e.g., code generation, result interpretation, presentation) to iteratively improve response quality.
-
Innovations:
- Creation of a design space integrating the three dimensions (semantic, rhetorical, and pragmatic), addressing core challenges in deriving insights from data.
- Adoption of a multi-agent approach, enabling collaborative and iterative improvements to LLM responses for greater completeness, reliability, and contextual relevance.
- Seamless integration of AI functionality into the popular Jupyter Notebook environment, reducing the need for tool-switching.
-
Implementation Steps:
- Extend the Jupyter Notebook layout with a dual-panel design for EDA and storytelling functionalities (left panel for the notebook, right sidebar for additional features).
- Use the multi-agent architecture to iteratively refine complex user requests, including multi-layered validation and response improvement.
- Provide different modes (single-agent & multi-agent) to balance response quality and latency.
- Enable users to control the analysis process through tooltips and clear result explanations, enhancing transparency and reparability.
Research Outcomes
-
Specific Outcomes:
- Development of "Jupybara," an AI assistant tool supporting the integration of EDA and data storytelling.
- Investigation and summary of key workflows and common pain points of data analysts, serving as design guidelines.
- Demonstration of the effectiveness of the two proposed strategies (design-space-aware prompting and multi-agent architecture) in enhancing LLM performance.
-
Advantages:
- Compared to Existing Tools:
- Compared to ChatGPT's analysis plugins, Jupybara excels in usability, interpretability, reparability, and integration.
- The multi-agent model improves the precision and quality of responses to complex data requests compared to simpler tools.
- Efficient Tool Integration: Eliminates frequent tool-switching by integrating into the familiar Jupyter environment.
- Targeted Optimization: Provides step-by-step guidance and iterative services for tasks such as insight extraction, result interpretation, and narrative strategy.
- Compared to Existing Tools:
-
Experimental or Evaluation Results:
- Based on experiments with 9 expert users experienced in data analysis, Jupybara was rated superior to ChatGPT plugins in terms of interpretability and applicability.
- The multi-agent mode (vs. single-agent mode) received higher scores in thematic precision, narrative coherence, and the delivery of high-quality actionable insights.
-
Limitations and Future Directions:
- Limitations:
- The sample size for user experiments was small, and most tasks were short-term, limiting the evaluation of the system's performance in large-scale, long-term projects.
- The system's response time is slower in multi-agent mode, potentially affecting user experience.
- Some auto-generated explanations (e.g., narrative prompts) lack richness and require further optimization.
- Future Research Directions:
- Explore automated switching mechanisms between single-agent and multi-agent modes to optimize efficiency.
- Integrate RAG (retrieval-augmented generation) methods to further improve the quality of domain-specific analyses.
- Expand support for other output formats such as slides and data videos to accommodate more diverse analysis presentation scenarios.
- Investigate how Jupybara can observe user-AI interactions to optimize collaboration efficiency for different user groups.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can semantic precision, narrative persuasiveness, and practical relevance be balanced to extract actionable insights in exploratory data analysis (EDA) and data storytelling?Category: Data Storytelling and Narrative Visualization NeedsSimilar questionsarrow_forward
- How can design-space-aware prompting optimize large language model (LLMs) performance in data analysis tasks?Category: Data Storytelling and Narrative Visualization NeedsSimilar questionsarrow_forward
- How can multi-agent architectures improve response quality and contextual relevance for complex data requests through division of labor?Category: Data Storytelling and Narrative Visualization NeedsSimilar questionsarrow_forward
Practical Problems
1- Data analysts struggle to efficiently extract actionable insights from complex data and communicate clear interpretations.Category: Data Storytelling and Narrative Visualization NeedsSimilar questionsarrow_forward
- 83%
GRAFS: Graphical Faceted Search System to Support Conceptual Understanding in Exploratory Search
IUI '24· Interactive Data Visualization +1
- 71%
Dango: A Mixed-Initiative Data Wrangling System using Large Language Model
CHI '25· Human-LLM Collaboration +2
- 71%
Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
PleaSQLarify: Visual Pragmatic Repair for Natural Language Database Querying
CHI '26· Human-LLM Collaboration +2
- 71%
SCSimulator: An Exploratory Visual Analytics Framework for Partner Selection in Supply Chains through LLM-driven Multi-Agent Simulation
IUI '26· Human-LLM Collaboration +2
- 71%
Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition
UIST '24· Human-LLM Collaboration +2
- 67%
ToonNote: Improving Communication in Computational Notebooks Using Interactive Data Comics
CHI '21· Interactive Data Visualization +1
- 67%
Visualizing Examples of Deep Neural Networks at Scale
CHI '21· Human-LLM Collaboration +1
- 67%
Tracing and Visualizing Human-ML/AI Collaborative Processes through Artifacts of Data Work
CHI '23· Human-LLM Collaboration +1
- 67%
"The Diagram is like Guardrails": Structuring GenAI-assisted Hypotheses Exploration with an Interactive Shared Representation
C&C '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)