VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
Authors
Research Background and Issues
-
What problems or challenges did the authors identify?
Automated planning tools have traditionally been designed for expert users, employing complex languages and logical methods that limit accessibility for general users. Furthermore, even advanced large language models (LLMs) face challenges in complex planning tasks, including unpredictability, susceptibility to "hallucinations" (generating false information), and a lack of user trust. -
Why is this issue important?
General users have complex planning needs in daily life, such as time management and project coordination. However, existing tools are not user-friendly and lack reliability, which limits their potential applications. Moreover, in critical domains such as healthcare or safety-critical tasks, these shortcomings could lead to negative safety implications. -
Research Motivation and Related Work
The authors propose enhancing the reliability and usability of these systems by integrating formal verification methods (e.g., model checking) with LLMs. By using model checking techniques to provide deterministic boundaries and incorporating more control mechanisms in user-system interactions, the approach aims to address inconsistencies and unmet constraints in planning.
Solution
-
What methods or solutions did the authors propose?
The authors developed a system called "VeriPlan," which combines large language models with formal verification methods (model checking) to address challenges in planning tasks. VeriPlan includes the following core features: a Rule Translator, Flexibility Sliders, and a Model Checker. -
What are the innovative aspects of this solution?
- Formal Verification: This is the first application of model checking technology to general users' planning tasks, verifying whether LLM outputs comply with user-defined constraints.
- User Involvement: By providing a Rule Translator and Flexibility Sliders, users can dynamically adjust the strictness and priority of rules, enhancing the effectiveness of human-computer interaction.
- Integration of LLMs and Formal Verification: The system transforms user inputs into logical constraints and employs external verification methods to optimize planning outputs.
-
What are the implementation steps and key technologies used?
- LLM Planning: Accepts natural language input from users and generates initial planning results.
- Rule Translator: Converts user input into formal logical expressions (LTL logic) and allows users to confirm the accuracy of the translation.
- Flexibility Sliders: Enables users to adjust the strictness of rules (soft constraints/hard constraints) via sliders.
- Model Checker: Uses the PRISM tool to verify whether the generated plan complies with user rules and provides feedback to both the user and the LLM.
- Iterative Optimization: The LLM iteratively generates plans based on feedback from the Model Checker until the rules are satisfied or iteration limits are reached.
Research Outcomes
-
What specific outcomes were achieved?
Through user studies (n=12), the effectiveness of VeriPlan was validated. Overall results showed that the VeriPlan system, which integrates the Rule Translator, Flexibility Sliders, and Model Checker, significantly improved the usability, output performance, and user trust in LLMs. -
What advantages does it have compared to existing solutions?
- Improved reliability in planning tasks: The Model Checker provides external validation to reduce errors in planning outputs.
- Enhanced user control: The Flexibility Sliders and Rule Translator allow users to dynamically adjust rule settings to meet personal preferences and situational needs.
- Innovation: The integration of formal verification methods with large language models enhances output transparency and user satisfaction. Users can actively define and optimize constraints.
-
What were the experimental or evaluation results?
In multiple task scenarios (e.g., patient navigation, chef optimization, and scheduling), the complete VeriPlan system (with all features) significantly outperformed variants that excluded the Model Checker (Condition 4) or lacked individual features (Conditions 2 and 3). User ratings showed significant improvements in system usability, satisfaction, and planning performance. -
Limitations and Future Directions
- Constraint Template Limitations: The current rule templates cover only a limited range of temporal constraints. Future research could expand to include more constraint types and allow users to define custom constraints.
- Model Framework Improvements: The PRISM tool restricts the types of logical expressions that can be used. Future work could explore new verification methods to enhance flexibility and expressiveness.
- Task Scope Expansion: Experiments were limited to a few everyday planning scenarios. Future research could test VeriPlan's performance in other domains (e.g., healthcare, manufacturing).
- Scalability Studies: Due to the small sample size, larger-scale user studies could be conducted to further validate the results.
Conclusion
VeriPlan opens a new pathway for automated planning by introducing formal verification techniques and user control, enhancing the reliability and usability of LLMs. This study demonstrates that combining model checking with user interaction effectively addresses various challenges in planning tasks. Future research could further expand functionality, improve flexibility, and deepen domain applications to meet the needs of users in complex, dynamic environments.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What problems do existing automated planning tools have in addressing ordinary users' everyday complex planning needs?Category: AI Trust Building and Reliability JudgmentSimilar questionsarrow_forward
- Can model checking techniques improve reliability and usability of LLMs in planning tasks?Category: AI Trust Building and Reliability JudgmentSimilar questionsarrow_forward
- How does combining user interaction with formal verification improve planning output quality and user trust?Category: AI Trust Building and Reliability JudgmentSimilar questionsarrow_forward
Practical Problems
1- Ordinary users struggle to use existing automated tools for complex planning with unreliable outputs.Category: AI Trust Building and Reliability JudgmentSimilar questionsarrow_forward
- 86%
How Do Analysts Understand and Verify AI-Assisted Data Analyses?
CHI '24· Human-LLM Collaboration +2
- 86%
PleaSQLarify: Visual Pragmatic Repair for Natural Language Database Querying
CHI '26· Human-LLM Collaboration +2
- 86%
SCSimulator: An Exploratory Visual Analytics Framework for Partner Selection in Supply Chains through LLM-driven Multi-Agent Simulation
IUI '26· Human-LLM Collaboration +2
- 86%
Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition
UIST '24· Human-LLM Collaboration +2
- 75%
Natural Language Dataset Generation Framework for Visualizations Powered by Large Language Models
CHI '24· Human-LLM Collaboration +2
- 75%
Lexara: A User-Centered Toolkit for Evaluating Large Language Models for Conversational Visual Analytics
CHI '26· Human-LLM Collaboration +2
- 75%
Steering Semantic Data Processing With DocWrangler
UIST '25· Human-LLM Collaboration +2
- 71%
iScore: Visual Analytics for Interpreting How Language Models Automatically Score Summaries
IUI '24· Explainable AI (XAI) +1
- 67%
DirectGPT: A Direct Manipulation Interface to Interact with Large Language Models
CHI '24· Human-LLM Collaboration +3
- 67%
Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models
IUI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)