Tisane: Authoring Statistical Models via Formal Reasoning from Conceptual and Data Relationships

Honorable Mention
Explainable AI (XAI)Computational Methods in HCIUniversity Professors & ResearchersStatisticians & Data Scientists

Document Title

Tisane: Authoring Statistical Models via Formal Reasoning from Conceptual and Data Relationships

Document Information

  • Subject Area: Human-Computer Interaction, Automation in Statistical Model Design and Inference
  • Keywords: Statistical analysis, Generalized Linear Models, Domain-Specific Language, Causal reasoning, Data science tools, Mixed-effects models, Seamless integration, Distributed data relationships, Data visualization, Statistical validity

Research Background and Problem

  • Identified Challenges:

    • Current statistical analysis tools fail to integrate domain knowledge, data collection processes, and model selection, making it difficult for users to avoid errors in statistical inference.
    • These tools often separate domain hypothesis reasoning from statistical model design, making it challenging for beginners to apply models effectively.
    • Incorrect statistical models can lead to theoretical errors, irreproducible conclusions, and unsupported public policies.
    • For instance, Generalized Linear Mixed-effects Models (GLMMs) perform well with hierarchical data, but these tools struggle to help users identify and correctly model hierarchical structures in data.
  • Significance:

    • Promotes scientific rigor and effectiveness in data analysis, especially for non-experts in statistics.
    • Reduces threats to the validity and external reliability of statistical conclusions.
  • Research Motivation and Related Work:

    • Current research tools like R and SPSS struggle to connect domain knowledge with data characteristics.
    • Existing end-to-end tools (e.g., Daggity and Tea) primarily focus on single-aspect analysis and lack comprehensive support for the entire process.

Solution

  • Solution Overview:

    • Introduced a hybrid interactive system called "Tisane," which supports a semi-automated modeling process for Generalized Linear Models (GLMs) and Generalized Linear Mixed-effects Models (GLMMs) by combining conceptual, data measurement relationships, and statistical modeling.
  • Innovations:

    • Developed a pioneering Study Design Specification Language (SDSL) that allows users to express relationships between variables at a high level of abstraction.
    • Built a graph-based Intermediate Representation (IR) to enable the system to infer candidate statistical models.
    • Designed an interactive compilation process that resolves ambiguities by querying users and outputs compliant models.
  • Implementation Steps and Key Techniques:

    1. Defining Variable Relationships: Users declare variables and annotate relationships (e.g., "causal") through the DSL.
    2. Pre-checking: Validates the conceptual correctness of model inputs to ensure logical consistency of causal relationships.
    3. Generating Candidate Models: Automatically infers main effects, interaction effects, and random effects based on the graph model.
    4. User Interaction for Disambiguation: A GUI guides users to confirm additional variables and effects, selecting appropriate distribution and link functions.
    5. Outputting Code and Records: Generates executable Python code and records user choices for auditing and reproducibility.

Research Outcomes

  • Specific Achievements:

    1. Provided a DSL tool that enables users to document data collection relationships and logically infer statistical models.
    2. Derived maximal random-effects structures supporting GLMMs, optimizing the external validity of statistical conclusions.
    3. Open-sourced the tool, facilitating seamless integration of quantitative domain knowledge and statistical modeling.
  • Advantages Over Existing Tools:

    • Compared to automated statistical tools (e.g., AutoML), Tisane emphasizes interpretability and coherence with domain knowledge.
    • Compared to languages like Tea, Tisane supports more complex hierarchical designs and interactive model inference.
    • Offers intuitive guidance for selecting random effects and family-link functions.
  • Experiments and Evaluation Results:

    • Case Studies: Tested in three real-world research scenarios (study planning, data analysis, and model iteration stages).
    • User feedback indicated that Tisane helped them focus more on research goals, enhanced explicit awareness of hypotheses, and reduced errors in previous analyses.
    • One researcher corrected errors in selecting between Generalized Linear Models and Mixed-effects Models after using Tisane.
  • Limitations and Future Directions:

    • Current model inference is limited to GLMs and GLMMs, without full coverage of causal relationship analysis.
    • Plans to develop features for simulation comparisons and sensitivity analysis to further enhance result robustness.
    • Aiming to implement the tool in other languages like R to expand its applicability.

Output Format

  • Final Generated Script: Provides executable scripts and residual plots to verify the appropriateness of family and link functions.
  • Experience Optimization Directions:
    • Add support for discipline-specific languages (e.g., psychology).
    • Improve GUI compatibility and support command-line execution in non-interactive environments.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/68888/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3501888
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
Honorable Mention
group
Authors
4 authors
sell
Subtopics
Explainable AI (XAI), Computational Methods in HCI
work
Professions
University Professors & Researchers, Statisticians & Data Scientists
article
Content Status
Full text indexed
hub
Related Papers
3 related papers