VeriPlan: Integrating Formal Verification and LLMs into End-User Planning

Human-LLM CollaborationExplainable AI (XAI)Interactive Data VisualizationSoftware Engineers & DevelopersUI/UX DesignersData Scientists & AnalystsAI/ML Researchers & Engineers

Research Background and Issues

  • What problems or challenges did the authors identify?
    Automated planning tools have traditionally been designed for expert users, employing complex languages and logical methods that limit accessibility for general users. Furthermore, even advanced large language models (LLMs) face challenges in complex planning tasks, including unpredictability, susceptibility to "hallucinations" (generating false information), and a lack of user trust.

  • Why is this issue important?
    General users have complex planning needs in daily life, such as time management and project coordination. However, existing tools are not user-friendly and lack reliability, which limits their potential applications. Moreover, in critical domains such as healthcare or safety-critical tasks, these shortcomings could lead to negative safety implications.

  • Research Motivation and Related Work
    The authors propose enhancing the reliability and usability of these systems by integrating formal verification methods (e.g., model checking) with LLMs. By using model checking techniques to provide deterministic boundaries and incorporating more control mechanisms in user-system interactions, the approach aims to address inconsistencies and unmet constraints in planning.

Solution

  • What methods or solutions did the authors propose?
    The authors developed a system called "VeriPlan," which combines large language models with formal verification methods (model checking) to address challenges in planning tasks. VeriPlan includes the following core features: a Rule Translator, Flexibility Sliders, and a Model Checker.

  • What are the innovative aspects of this solution?

    1. Formal Verification: This is the first application of model checking technology to general users' planning tasks, verifying whether LLM outputs comply with user-defined constraints.
    2. User Involvement: By providing a Rule Translator and Flexibility Sliders, users can dynamically adjust the strictness and priority of rules, enhancing the effectiveness of human-computer interaction.
    3. Integration of LLMs and Formal Verification: The system transforms user inputs into logical constraints and employs external verification methods to optimize planning outputs.
  • What are the implementation steps and key technologies used?

    1. LLM Planning: Accepts natural language input from users and generates initial planning results.
    2. Rule Translator: Converts user input into formal logical expressions (LTL logic) and allows users to confirm the accuracy of the translation.
    3. Flexibility Sliders: Enables users to adjust the strictness of rules (soft constraints/hard constraints) via sliders.
    4. Model Checker: Uses the PRISM tool to verify whether the generated plan complies with user rules and provides feedback to both the user and the LLM.
    5. Iterative Optimization: The LLM iteratively generates plans based on feedback from the Model Checker until the rules are satisfied or iteration limits are reached.

Research Outcomes

  • What specific outcomes were achieved?
    Through user studies (n=12), the effectiveness of VeriPlan was validated. Overall results showed that the VeriPlan system, which integrates the Rule Translator, Flexibility Sliders, and Model Checker, significantly improved the usability, output performance, and user trust in LLMs.

  • What advantages does it have compared to existing solutions?

    1. Improved reliability in planning tasks: The Model Checker provides external validation to reduce errors in planning outputs.
    2. Enhanced user control: The Flexibility Sliders and Rule Translator allow users to dynamically adjust rule settings to meet personal preferences and situational needs.
    3. Innovation: The integration of formal verification methods with large language models enhances output transparency and user satisfaction. Users can actively define and optimize constraints.
  • What were the experimental or evaluation results?
    In multiple task scenarios (e.g., patient navigation, chef optimization, and scheduling), the complete VeriPlan system (with all features) significantly outperformed variants that excluded the Model Checker (Condition 4) or lacked individual features (Conditions 2 and 3). User ratings showed significant improvements in system usability, satisfaction, and planning performance.

  • Limitations and Future Directions

    1. Constraint Template Limitations: The current rule templates cover only a limited range of temporal constraints. Future research could expand to include more constraint types and allow users to define custom constraints.
    2. Model Framework Improvements: The PRISM tool restricts the types of logical expressions that can be used. Future work could explore new verification methods to enhance flexibility and expressiveness.
    3. Task Scope Expansion: Experiments were limited to a few everyday planning scenarios. Future research could test VeriPlan's performance in other domains (e.g., healthcare, manufacturing).
    4. Scalability Studies: Due to the small sample size, larger-scale user studies could be conducted to further validate the results.

Conclusion

VeriPlan opens a new pathway for automated planning by introducing formal verification techniques and user control, enhancing the reliability and usability of LLMs. This study demonstrates that combining model checking with user interaction effectively addresses various challenges in planning tasks. Future research could further expand functionality, improve flexibility, and deepen domain applications to meet the needs of users in complex, dynamic environments.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188642/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714113
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), Interactive Data Visualization
work
Professions
Software Engineers & Developers, UI/UX Designers, Data Scientists & Analysts, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers