Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication
Authors
Document Title
Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication
Document Information
- Subject Area: Human-Computer Interaction (HCI), Language Model Applications, AI-Assisted Writing
- Keywords: AI-Assisted Writing, Large Language Models, Information Exchange, Political Communication, GPT-3, Government Staff, User Experience, Automated Suggestions, Human Collaboration, Technology Ethics
Research Background and Problem
- Identified Challenges: Traditional writing assistance systems often focus on word or phrase suggestions, whereas current large language models (e.g., GPT-3) have the capability to generate longer text suggestions. While these models hold potential in high-demand communication environments (e.g., legislative staff responding to large volumes of constituent emails), their impact on user experience and outcome quality remains unclear.
- Significance: High-quality communication is critical for building trust, especially in political contexts. AI-generated suggestions may improve efficiency but could also undermine trust between constituents and officials.
- Research Motivation and Related Work: Existing studies largely focus on phrase-level suggestions from language models, with limited exploration of automatically generated full draft responses. This study aims to investigate the trade-offs and applicable scenarios for sentence-level and message-level suggestions in human-led communication. Related work includes research on language model ethics, text generation quality, and the technological impact on political communication.
Solution
- Proposed Solution: Development of an online platform called “Dispatch,” enabling users to simulate legislative staff responding to constituent emails using sentence-level and message-level suggestions, and designing experiments to compare the impact of these suggestions on the writing process.
- Innovations:
- Utilizing GPT-3 to generate sentence-level suggestions (focused on specific email content or extending the current draft) and message-level suggestions (producing full response drafts).
- Systematic comparison of the effects of these suggestions on response efficiency, user satisfaction, text quality, and writing autonomy.
- Implementation Steps and Key Technologies:
- Develop the Dispatch platform to support sentence-level and message-level suggestions (with interactive visual demonstrations).
- Use the GPT-3 model (text-davinci-002) to generate suggestion content.
- Select content from public letters as constituent email samples and design experimental conditions: no suggestions, sentence-level suggestions, message-level suggestions.
- Recruit 120 participants to simulate staff roles, complete tasks, and track their behaviors, collecting data on task completion time, editing actions, and text modifications.
Research Findings
- Specific Results:
- Message-level suggestions enabled participants to complete tasks faster, with an average time of 8.53 minutes, significantly lower than sentence-level suggestions (15.77 minutes) and no suggestions (16.4 minutes).
- Responses generated with message-level suggestions were rated as the most helpful by third-party evaluators, outperforming sentence-level suggestions and fully manual responses.
- Sentence-level suggestions allowed participants to retain more autonomy in writing; approximately 49.4% of the final response content was manually added, compared to 24.25% in message-level suggestions.
- Advantages Over Existing Solutions:
- Message-level suggestions significantly improved efficiency and produced naturally fluent response drafts.
- Sentence-level suggestions enhanced user creativity and engagement, favoring human-led communication processes.
- Experimental or Evaluation Results:
- Responses generated with message-level suggestions demonstrated superior lexical diversity and lower grammar error rates, with an error rate of 0.158/word, significantly lower than fully manual responses (0.176/word).
- Message-level suggestions were perceived as more helpful than unassisted manual responses, even when AI involvement was disclosed, outperforming sentence-level suggestions.
- Limitations and Future Directions:
- Sentence-level suggestions are constrained by context windows, affecting generation quality.
- In political communication scenarios, AI usage must be cautious to avoid further erosion of trust between constituents and officials.
- Future research could explore diversified recommendation options, enhanced long-text generation capabilities, and more personalized support for communication objectives.
Design Implications
- The choice of AI suggestion units should flexibly adapt to task requirements, as different communication contexts (e.g., personal birthday messages versus professional replies) may necessitate varying suggestion formats.
- Providing multiple alternative suggestions can enhance flexibility but should avoid excessive redundancy while ensuring diversity among options.
- Attention should be given to the quality and customization of model-generated results, such as fine-tuning generation modes to match specific tones or length requirements.
- Beyond text suggestions, additional support features could be introduced, such as highlighting key points, providing contextual background information, or tracking past responses.
- Transparency about AI involvement may influence user trust in the responses, necessitating a balance between privacy and transparency.
Conclusion
This study demonstrates the distinct roles of different suggestion types in AI-mediated writing through experimental evidence, uncovering how communication assistance technologies impact efficiency, autonomy, and interaction quality. The findings aim to inform the design of communication assistance systems tailored to specific contexts and call for further exploration of the application potential of large language models across domains and languages.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do sentence-level and message-level suggestions differ in effectiveness for AI-assisted writing?Category: Public Perception of AI and Algorithmic AccountabilitySimilar questionsarrow_forward
- Which suggestion form (sentence-level or message-level) better improves efficiency, user satisfaction, and interaction quality?Category: Public Perception of AI and Algorithmic AccountabilitySimilar questionsarrow_forward
- How can sentence-level suggestions provide effective support while preserving user autonomy?Category: Public Perception of AI and Algorithmic AccountabilitySimilar questionsarrow_forward
Practical Problems
1- Government staff respond to large volumes of email inefficiently and easily lose public trust.Category: Public Perception of AI and Algorithmic AccountabilitySimilar questionsarrow_forward
- 100%
VAL: Interactive Task Learning with GPT Dialog Parsing
CHI '24· Human-LLM Collaboration +1
- 100%
Need Help? Designing Proactive AI Assistants for Programming
CHI '25· Human-LLM Collaboration +1
- 80%
Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making
CHI '23· Human-LLM Collaboration +1
- 80%
Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts
CHI '23· Human-LLM Collaboration +1
- 80%
Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming
CHI '24· Human-LLM Collaboration +2
- 80%
Automatic Macro Mining from Interaction Traces at Scale
CHI '24· Human-LLM Collaboration +1
- 80%
"If the Machine Is As Good As Me, Then What Use Am I?" – How the Use of ChatGPT Changes Young Professionals' Perception of Productivity and Accomplishment
CHI '24· Human-LLM Collaboration +1
- 80%
Interactive Debugging and Steering of Multi-Agent AI Systems
CHI '25· Human-LLM Collaboration +2
- 80%
GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 80%
New Enactions of Expertise: Software Engineers’ Evaluation and Demonstration of Coding Expertise with AI Coding Assistants
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)