TurnStyle: A Framework for Analyzing Human Conversational Behaviors to Predict Success in LLM-Assisted Tasks

Honorable Mention
Human-LLM CollaborationAI-Assisted Decision-Making & AutomationUser Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingUniversity Professors & ResearchersSoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI Researchers

Paper Title

TurnStyle: A Framework for Analyzing Human Conversational Behaviors to Predict Success in LLM-Assisted Tasks

Publication Info

  • Topic area: Human–LLM interaction and task success prediction
  • Keywords: Human–AI collaboration, large language models, conversational analysis, task success, taxonomy, collaborative learning, sequence modeling, predictive modeling, task management, information request

Background and Problem

  • Problem / challenge: Existing studies on human–LLM collaboration focus on static, domain-specific prompting strategies that do not generalize across contexts or adapt to evolving LLM capabilities. There is a lack of task-agnostic, turn-level frameworks to characterize human conversational behaviors that predict task success.
  • Significance: Understanding and improving human–LLM collaboration is critical as LLMs are increasingly used in education and professional settings for cognitively demanding tasks. A robust framework can guide training, interface design, and adaptive support tools.
  • Motivation and related work: Prior work includes domain-specific taxonomies (e.g., customer service chatbots) and broad conversational frameworks (e.g., IBM Natural Conversation Framework), but these lack task-agnostic applicability or turn-level granularity. Existing datasets often lack multi-turn dialogues or objective task outcomes. TurnStyle builds on collaborative learning theory and conversational dynamics to address these gaps.

Solution

  • Proposed approach: TurnStyle, a turn-level, task-agnostic taxonomy for analyzing human conversational behaviors in human–LLM interactions, linking these behaviors to task success.
  • Novelty:
    1. Development of a domain-agnostic taxonomy adapted from collaborative learning theory, with categories tailored for human–LLM asymmetry.
    2. Empirical validation across three datasets (StudyChat, DevGPT, workplace reskilling trial) with outcome labels, identifying cross-domain conversational patterns linked to success.
    3. A coaching playbook of five trainable conversational behaviors for effective human–LLM collaboration.
    4. Open-source resources for replicating and extending TurnStyle, including annotation prompts, scripts, and annotated datasets.
  • Procedure and key techniques:
    • Iterative development of the taxonomy through pilot annotations and refinement.
    • Annotation of 26,335 human turns across 3,365 conversations using a hybrid LLM–human pipeline.
    • Sequence and predictive analyses, including Markov modeling, Hidden Markov Models (HMMs), and logistic regression, to link conversational behaviors to task outcomes.

Results

  • Concrete findings:
    • Spending excessive time in inquiry loops (e.g., repeated domain-level questions) predicts lower success (OR = 0.714, CI [0.533, 0.957], p = 0.024).
    • Transitions from task definition to agreement (e.g., validating scope) are enriched in successful conversations (OR = 1.484, CI [1.115, 1.976]).
    • Early focus on formatting is detrimental to success (OR = 0.584, CI [0.365, 0.933]).
    • Successful conversations exhibit higher entropy rates (+0.282 bits/turn) and greater diversity in conversational styles.
  • Advantage over baselines:
    • TurnStyle identifies task-agnostic conversational patterns that generalize across datasets, unlike domain-specific or static prompting strategies.
    • Predictive models using TurnStyle features outperform simple length-based or static feature models.
  • Experiments / evaluation:
    • Datasets: StudyChat (1,234 conversations), DevGPT (1,655 conversations), workplace reskilling trial (476 conversations).
    • Metrics: Success defined by grades, issue/PR resolution, or benchmark scores.
    • Techniques: Local and phase-aware transition enrichment, Markov dynamics, HMMs, and logistic regression.
  • Limitations and future work:
    • Data limitations: External resources used during tasks are not captured in conversational logs.
    • Annotation biases: LLM-based annotations may reflect model-specific priors.
    • Generalizability: Framework validated in STEM domains; further work needed for creative or social tasks.

Summary

TurnStyle introduces a task-agnostic, turn-level taxonomy for analyzing human conversational behaviors in LLM-assisted tasks, linking these behaviors to measurable success outcomes. Validated across three datasets, TurnStyle identifies actionable conversational patterns, such as transitioning from inquiry to task management and validating scope, that predict success. The framework is resilient to evolving LLM capabilities and provides open resources for replication and extension. TurnStyle has implications for training, interface design, and real-time adaptive support, with potential applications beyond STEM into creative and organizational domains.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222870/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790459
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
3 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, User Research Methods (Interviews, Surveys, Observation), Prototyping & User Testing
work
Professions
University Professors & Researchers, Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers