TurnStyle: A Framework for Analyzing Human Conversational Behaviors to Predict Success in LLM-Assisted Tasks
Honorable MentionAuthors
Paper Title
TurnStyle: A Framework for Analyzing Human Conversational Behaviors to Predict Success in LLM-Assisted Tasks
Publication Info
- Topic area: Human–LLM interaction and task success prediction
- Keywords: Human–AI collaboration, large language models, conversational analysis, task success, taxonomy, collaborative learning, sequence modeling, predictive modeling, task management, information request
Background and Problem
- Problem / challenge: Existing studies on human–LLM collaboration focus on static, domain-specific prompting strategies that do not generalize across contexts or adapt to evolving LLM capabilities. There is a lack of task-agnostic, turn-level frameworks to characterize human conversational behaviors that predict task success.
- Significance: Understanding and improving human–LLM collaboration is critical as LLMs are increasingly used in education and professional settings for cognitively demanding tasks. A robust framework can guide training, interface design, and adaptive support tools.
- Motivation and related work: Prior work includes domain-specific taxonomies (e.g., customer service chatbots) and broad conversational frameworks (e.g., IBM Natural Conversation Framework), but these lack task-agnostic applicability or turn-level granularity. Existing datasets often lack multi-turn dialogues or objective task outcomes. TurnStyle builds on collaborative learning theory and conversational dynamics to address these gaps.
Solution
- Proposed approach: TurnStyle, a turn-level, task-agnostic taxonomy for analyzing human conversational behaviors in human–LLM interactions, linking these behaviors to task success.
- Novelty:
- Development of a domain-agnostic taxonomy adapted from collaborative learning theory, with categories tailored for human–LLM asymmetry.
- Empirical validation across three datasets (StudyChat, DevGPT, workplace reskilling trial) with outcome labels, identifying cross-domain conversational patterns linked to success.
- A coaching playbook of five trainable conversational behaviors for effective human–LLM collaboration.
- Open-source resources for replicating and extending TurnStyle, including annotation prompts, scripts, and annotated datasets.
- Procedure and key techniques:
- Iterative development of the taxonomy through pilot annotations and refinement.
- Annotation of 26,335 human turns across 3,365 conversations using a hybrid LLM–human pipeline.
- Sequence and predictive analyses, including Markov modeling, Hidden Markov Models (HMMs), and logistic regression, to link conversational behaviors to task outcomes.
Results
- Concrete findings:
- Spending excessive time in inquiry loops (e.g., repeated domain-level questions) predicts lower success (OR = 0.714, CI [0.533, 0.957], p = 0.024).
- Transitions from task definition to agreement (e.g., validating scope) are enriched in successful conversations (OR = 1.484, CI [1.115, 1.976]).
- Early focus on formatting is detrimental to success (OR = 0.584, CI [0.365, 0.933]).
- Successful conversations exhibit higher entropy rates (+0.282 bits/turn) and greater diversity in conversational styles.
- Advantage over baselines:
- TurnStyle identifies task-agnostic conversational patterns that generalize across datasets, unlike domain-specific or static prompting strategies.
- Predictive models using TurnStyle features outperform simple length-based or static feature models.
- Experiments / evaluation:
- Datasets: StudyChat (1,234 conversations), DevGPT (1,655 conversations), workplace reskilling trial (476 conversations).
- Metrics: Success defined by grades, issue/PR resolution, or benchmark scores.
- Techniques: Local and phase-aware transition enrichment, Markov dynamics, HMMs, and logistic regression.
- Limitations and future work:
- Data limitations: External resources used during tasks are not captured in conversational logs.
- Annotation biases: LLM-based annotations may reflect model-specific priors.
- Generalizability: Framework validated in STEM domains; further work needed for creative or social tasks.
Summary
TurnStyle introduces a task-agnostic, turn-level taxonomy for analyzing human conversational behaviors in LLM-assisted tasks, linking these behaviors to measurable success outcomes. Validated across three datasets, TurnStyle identifies actionable conversational patterns, such as transitioning from inquiry to task management and validating scope, that predict success. The framework is resilient to evolving LLM capabilities and provides open resources for replication and extension. TurnStyle has implications for training, interface design, and real-time adaptive support, with potential applications beyond STEM into creative and organizational domains.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
"Shall We Dig Deeper?": Designing and Evaluating Strategies for LLM Agents to Advance Knowledge Co-Construction in Asynchronous Online Discussions
CHI '26· Human-LLM Collaboration +3
- 88%
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
CHI '26· Human-LLM Collaboration +3
- 88%
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
CHI '26· Human-LLM Collaboration +3
- 88%
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
CHI '26· Human-LLM Collaboration +3
- 88%
Live in the Loop: Rapid Run-time Feedback for Prompts
CHI '26· Human-LLM Collaboration +3
- 78%
Interview-Informed Generative Agents for Product Discovery: A Validation Study
CHI '26· Human-LLM Collaboration +3
- 78%
CHOIR: A Chatbot-mediated Organizational Memory Leveraging Communication in University Research Labs
CHI '26· Human-LLM Collaboration +3
- 78%
Just-In-Time Objectives: A General Approach for Specialized AI Interactions
CHI '26· Human-LLM Collaboration +3
- 75%
Computer Science Achievement and Writing Skills Predict Vibe Coding Proficiency
CHI '26· Human-LLM Collaboration +2
- 75%
InterFlow: Designing Unobtrusive AI to Empower Interviewers in Semi-Structured Interviews
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)