From Conversation to Human-AI Common Ground: Extracting Cognitive Workflows for Reuse in Sense-making Tasks
Authors
Paper Title
From Conversation to Human-AI Common Ground: Extracting Cognitive Workflows for Reuse in Sense-making Tasks
Publication Info
- Topic area: Human-AI collaboration in sense-making tasks with a focus on workflow reuse.
- Keywords: conversational AI, sense-making, cognitive workflows, common ground, intent expression, workflow reuse, knowledge work, schema extraction, human-AI collaboration, adaptive reasoning.
Background and Problem
- Problem / challenge: Current conversational AI systems fail to maintain a coherent understanding of evolving task structures, leading to repetitive context reconstruction and misalignment in multi-turn sense-making tasks.
- Significance: Addressing these limitations is crucial for improving efficiency and reducing cognitive load in knowledge-intensive tasks such as market analysis, research synthesis, and content creation.
- Motivation and related work: Existing tools like prompt engineering, memory systems, and workflow automation either lack adaptability for evolving tasks or require upfront specification, which is impractical for tacit and iterative reasoning. This paper builds on cognitive task analysis and workflow provenance research to address these gaps.
Solution
- Proposed approach: ThinkFlow, a system that transforms multi-turn human-AI conversations into reasoning-aware cognitive workflows, enabling dynamic common ground and reusable task structures.
- Novelty:
- Introduction of a reasoning-aware workflow schema that captures goals, decision criteria, and dependencies from conversations.
- A bidirectional chat-canvas interface for visualizing, editing, and grounding workflows.
- Empirical evidence demonstrating improved AI alignment and adaptive reuse through shared reasoning representations.
- Procedure and key techniques:
- AI extracts a structured workflow schema in real-time from conversations.
- Users inspect and refine the schema through direct manipulation or natural language feedback.
- The system enables flexible reuse of workflows at multiple granularities (components, phases, or entire workflows) across contexts.
Results
- Concrete findings:
- Schema fidelity: Expert raters scored the extracted schemas highly (mean = 6.29/7), confirming accurate representation of reasoning.
- Reuse quality: The schema condition outperformed baselines (chat history and task-only) in structure, accuracy, and adaptation, with significant improvements in final deliverables.
- Advantage over baselines: The schema condition consistently provided better reasoning preservation, decision carryover, and contextual adaptation compared to chat history and task-only conditions.
- Experiments / evaluation:
- Study 1: Expert evaluation of schema fidelity and reuse quality using 8 simulated conversations across 4 domains.
- Study 2: User study with 8 participants performing real-world tasks, demonstrating ThinkFlow's support for reasoning visibility, intent expression, and flexible reuse.
- Limitations and future work:
- Current schema extraction is resource-intensive and may not suit single-turn or casual tasks.
- Limited evaluation of cross-domain reuse and very long sessions.
- Need for mixed-initiative control over schema extraction and reuse.
Summary
ThinkFlow addresses the challenge of maintaining dynamic common ground in human-AI sense-making tasks by introducing a reasoning-aware workflow schema and a bidirectional chat-canvas interface. The system externalizes and structures reasoning, enabling users to inspect, refine, and reuse workflows across contexts. Empirical studies demonstrate that ThinkFlow improves reasoning alignment, output quality, and adaptive reuse compared to baseline approaches. While promising for complex, evolving tasks, future work will focus on optimizing schema extraction, supporting cross-domain reuse, and balancing automation with user control.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)