Cocoa: Co-Planning and Co-Execution with AI Agents
Best PaperAuthors
Paper Title
Cocoa: Co-Planning and Co-Execution with AI Agents
Publication Info
- Topic area: Human-AI collaboration in scientific research workflows
- Keywords: AI agents, human-AI collaboration, co-planning, co-execution, scientific research, computational notebooks, interactive systems, large language models, task delegation, document editing
Background and Problem
- Problem / challenge: Existing AI systems for research often rigidly separate planning and execution or rely on fixed collaborative workflows, limiting flexibility and user control. These systems lack mechanisms for fluid human-AI collaboration and fail to support iterative adjustments during task execution.
- Significance: Addressing these limitations can improve the usability, steerability, and effectiveness of AI agents in complex, knowledge-intensive domains like scientific research.
- Motivation and related work: Prior systems have explored agent-guided and user-guided workflows but lack flexible delegation of agency. Fully autonomous systems like ReAct interleave planning and execution but do not involve human collaboration. Computational notebooks offer inspiration for iterative workflows but have not been adapted for human-AI collaboration in research.
Solution
- Proposed approach: Cocoa, an interactive system that integrates AI agents into a document editor to enable interleaved co-planning and co-execution for scientific research tasks.
- Novelty:
- Introduction of interleaved co-planning and co-execution, allowing users and AI agents to collaboratively plan and execute tasks iteratively.
- Flexible delegation of tasks between users and AI agents, enabling context-dependent adjustments.
- Integration of AI agents into a document editing environment, inspired by computational notebooks, to support intuitive and interactive workflows.
- Procedure and key techniques:
- Users invoke the agent to propose plans, which are editable and interactive within the document.
- Tasks within the plan can be assigned to the user or the agent, with options for stepwise or continuous execution.
- Users can edit outputs, re-plan steps, and interleave planning and execution dynamically.
- Outputs are presented in an interactive sidebar or as a final summary panel in the document.
Results
- Concrete findings:
- In a lab study (n = 16), Cocoa improved agent steerability (median rating of 4 vs. 3 for a chat baseline, p = 0.005) without sacrificing ease of use.
- In a field deployment (n = 7), users leveraged Cocoa for literature synthesis and early-stage project planning, assigning high-level tasks to themselves and iterative tasks to the agent.
- Advantage over baselines:
- Cocoa reduced passive output inspection time (33.6% vs. 54.1% in the baseline) and increased active engagement through co-execution (15.2% vs. 3.1%, p < 0.001).
- Participants found Cocoa’s structured, editable plans more effective for steering the agent compared to the chat interface.
- Experiments / evaluation:
- Lab study: Within-subjects design comparing Cocoa to a chat-based baseline for literature-augmented tasks.
- Deployment study: 7-day field use of Cocoa for real-world research projects, tracking task delegation and iterative workflows.
- Limitations and future work:
- Cocoa’s linear document format limits plan forking and parallel iteration; future work could explore node-based canvases.
- The system lacks multimodal capabilities and code execution support.
- Studies were limited to researchers in computer science and adjacent fields; broader domains need exploration.
Summary
Cocoa is an interactive system that enables scientific researchers to collaboratively plan and execute tasks with AI agents in a document editing environment. By interleaving co-planning and co-execution, Cocoa allows users to iteratively refine plans and outputs, providing greater flexibility and control compared to traditional chat-based interfaces. Lab and field studies demonstrated that Cocoa improves agent steerability and supports dynamic human-AI collaboration without compromising ease of use. Future work could expand Cocoa’s capabilities to support multimodal inputs, hierarchical plans, and broader research domains.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Modelling Experts' Sampling Strategy to Balance Multiple Objectives During Scientific Explorations
HRI '24· Human-LLM Collaboration +2
- 71%
Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
IUI '26· Human-LLM Collaboration +2
- 67%
Formulating or Fixating: Effects of Examples on Problem Solving Vary as a Function of Example Presentation Interface Design
CHI '24· Prototyping & User Testing +1
- 67%
Evaluating Large Language Models on Academic Literature Understanding and Review: An Empirical Study among Early-stage Scholars
CHI '24· Human-LLM Collaboration +1
- 63%
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
CHI '26· Human-LLM Collaboration +3
- 63%
Criticality: Scaffolding Decision-Making with Interactive Critical Thinking and Evidence-Based Reasoning Traces
IUI '26· Human-LLM Collaboration +3
- 63%
Key Considerations for Domain Expert Involvement in LLM Design and Evaluation: An Ethnographic Study
IUI '26· Human-LLM Collaboration +3
- 63%
Marcelle: Composing Interactive Machine Learning Workflows and Interfaces
UIST '21· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)