Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models
Authors
Title of the Paper
Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models
Paper Information
- Subject Area: Human-Computer Interaction (HCI), Artificial Intelligence, and Collaborative Writing
- Keywords: AI-generated text, collaborative writing, human-AI collaboration, generative language models, writing assistants
Research Background and Problem
- Identified Issues or Challenges: With the advancements in generative language models (e.g., GPT-3), their potential in collaborative writing has garnered significant attention. However, there is still a lack of comprehensive understanding of how varying levels of AI support influence the human writing process and outcomes.
- Significance: AI writing tools hold immense potential to enhance creativity and facilitate expression. However, their design must ensure human agency to maintain user satisfaction and a sense of ownership over the text. Research on balancing AI support with human autonomy is crucial for the future design of writing tools.
- Research Motivation and Related Work:
- Traditional writing support tools typically offer simple completion functions but fail to effectively assist with complex writing needs.
- Previous studies have shown the potential of generative AI to improve efficiency in areas such as drafting help requests, storytelling, and academic writing. However, the impact of different levels of AI support on user experience remains underexplored.
- Related works referenced include studies on WordCraft and CoAuthor, which analyze the capabilities and applicability of generative AI in collaborative writing.
Solution
- Proposed Methods or Solutions:
- Develop an experimental platform using a generative language model (GPT-3) to simulate human-AI collaborative writing scenarios.
- Design three levels of AI support: no AI support (baseline condition), single-sentence suggestions (low-level support), and full-paragraph suggestions (high-level support). Analyze their effects on writing quality and user experience.
- Conduct classroom experiments with 131 participants to collect real-world writing data and subjective feedback.
- Innovations:
- Developed a multi-level scaffolding framework for AI-assisted writing.
- Used a U-shaped relationship model to analyze the impact of low and high levels of support on writing quality, user satisfaction, and productivity.
- Quantified how technical proficiency and writing skills influence outcomes in human-AI collaboration.
- Implementation Steps and Key Technologies:
- Employed a Latin Square Design in the experiments to balance conditions and minimize order effects.
- Utilized multiple measurement methods, including human ratings, automated tools (e.g., TAACO and TAALES), and NASA Task Load Index.
- Collected behavioral data on user editing habits and evaluated user satisfaction and sense of ownership in AI collaboration.
Research Findings
- Specific Findings:
- Provided evidence on how varying levels of scaffolding influence writing quality: low-level scaffolding (single-sentence suggestions) may disrupt text coherence, while high-level scaffolding (paragraph suggestions) significantly improves writing quality.
- High-level scaffolding notably enhanced the writing performance of non-expert and less tech-savvy users.
- While AI suggestions increased productivity (words per unit time), they also led to a decrease in user satisfaction and sense of ownership over the text.
- No significant impact on cognitive load was observed.
- Advantages: Compared to the no-support condition, high-level scaffolding offered a more coherent article structure, significantly reduced the amount of content participants needed to edit, and improved quality scores.
- Experimental or Evaluation Results:
- The U-shaped relationship showed that low-level scaffolding reduced text output quality, while high-level scaffolding significantly improved it.
- Data indicated increased user reliance on paragraph-level AI suggestions, which may reduce writing autonomy.
- Limitations and Future Directions:
- Did not explore complex scaffolding systems that dynamically adjust support levels.
- The experiment was limited to English-speaking users and did not include multilingual or low-writing-skill groups.
- The experimental scenario was confined to short-term testing, without observing the long-term impact of AI tools on writing habits.
- Future research could expand to more diverse writing tasks (narrative, technical writing, etc.) and real-world writing contexts (academic, journalistic writing, etc.).
Through this study, the authors emphasize the necessity of designing user-centered AI writing tools. These tools should dynamically adjust support levels based on diverse user needs to better foster creativity and productivity while maintaining user satisfaction and engagement.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do different levels of AI support affect human writing quality?Category: Writing Collaboration, Summarization, and Text SuggestionsSimilar questionsarrow_forward
- What impact do AI sentence-level suggestions versus full-paragraph suggestions have on user experience?Category: Writing Collaboration, Summarization, and Text SuggestionsSimilar questionsarrow_forward
- How do technical proficiency and writing ability affect outcomes in human-AI writing collaboration?Category: Writing Collaboration, Summarization, and Text SuggestionsSimilar questionsarrow_forward
Practical Problems
1- AI-assisted writing tools may weaken users' sense of text ownership and satisfaction.Category: Writing Collaboration, Summarization, and Text SuggestionsSimilar questionsarrow_forward
- 83%
The HaLLMark Effect: Supporting Provenance and Transparent Use of Large Language Models in Writing with Interactive Visualization
CHI '24· Human-LLM Collaboration +2
- 67%
PhraseFlow: Designs and Empirical Studies of Phrase-Level Input
CHI '21· Generative AI (Text, Image, Music, Video) +2
- 67%
Choice Over Control: How Users Write with Large Language Models using Diegetic and Non-Diegetic Prompting
CHI '23· Human-LLM Collaboration +1
- 67%
Art or Artifice? Large Language Models and the False Promise of Creativity
CHI '24· Human-LLM Collaboration +1
- 67%
Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits
CHI '25· Human-LLM Collaboration +1
- 67%
SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 60%
Metaphoria: An Algorithmic Companion for Metaphor Creation
CHI '19· Human-LLM Collaboration +1
- 60%
Tap&Say: Touch Location-Informed Large Language Model for Multimodal Text Correction on Smartphones
CHI '25· Human-LLM Collaboration
- 60%
Creativity Support in the Age of Large Language Models: An Empirical Study Involving Professional Writers
C&C '24· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)