Choice Over Control: How Users Write with Large Language Models using Diegetic and Non-Diegetic Prompting
Authors
Document Title
Choice Over Control: How Users Write with Large Language Models using Diegetic and Non-Diegetic Prompting
Document Information
- Authors: Hai Dang, Sven Goller, Florian Lehmann, Daniel Buschek
- Year of Publication: 2023
- Conference: CHI ’23: ACM Conference on Human Factors in Computing Systems
- Field: Human-Computer Interaction (HCI), Generative Natural Language Processing (NLP)
- Keywords: Large Language Models (LLM), Collaborative Writing Systems, Human-AI Collaboration, User-Driven Natural Language Generation, Prompt Engineering
Research Background and Problem
-
Identified Problems or Challenges:
- Existing prompt designs primarily focus on non-diegetic prompts, neglecting the role of diegetic prompts.
- It remains unclear how users choose or control suggestions generated by language models, particularly in their use of diegetic and non-diegetic prompts.
- Non-diegetic prompts increase users' cognitive load, and their impact on the writing experience requires further exploration.
-
Significance: User behavior and needs during collaboration with generative models directly influence the efficiency and quality of interaction design. Exploring prompt types can reveal the optimal integration of generative models into users' writing workflows, aiding in the development of more efficient and intuitive human-AI collaboration tools.
-
Research Motivation: By distinguishing between diegetic and non-diegetic prompts, this study aims to elevate the design of user interfaces powered by large language models (LLMs) to a new perspective, thereby optimizing user experience.
Solution
-
Proposed Approach:
- Introduce a conceptual framework distinguishing diegetic prompts (text content from the user's writing) and non-diegetic prompts (explicit instructions provided by the user).
- Design and test four user interface (UI) variants, including single and multiple suggestions, as well as the option to use non-diegetic prompts, to explore user writing behavior patterns in different scenarios.
-
Innovations:
- Integration of diegetic and non-diegetic prompts into a unified user interface, enabling users to freely choose their preferred prompting method during the writing process.
- Development of a granular user behavior analysis framework to evaluate how users receive and guide LLM-generated suggestions.
-
Implementation Steps and Key Techniques:
- Experimental Design: Conducted a remote user study with 129 participants, asking them to complete five topic-based writing tasks while testing four user interface variants with different prompt types.
- Technical Implementation: Utilized OpenAI's GPT-3 (text-davinci-edit-001) as the language model, built the front-end with ReactJS, and handled back-end requests using FastAPI.
- Data Collection:
- Logged interaction behaviors, including suggestion trigger times and acceptance rates.
- Conducted surveys to gather users' subjective evaluations of the interfaces.
- Assessed the quality of user-generated text, including grammatical errors and thematic coherence.
Research Findings
-
Key Results:
- Users preferred selecting multiple suggestions over providing explicit instructions via non-diegetic prompts.
- While non-diegetic prompts allowed for more precise control over text generation, they could disrupt the writing flow during short writing tasks.
- Users strategically adjusted the trigger positions for suggestions: single suggestions were more often triggered mid-sentence or after specific verbs, while multiple suggestions were more frequently triggered at the beginning of sentences.
-
Advantages and Contributions:
- The classification method for prompt interaction design (diegetic vs. non-diegetic) offers a new perspective for studying user collaboration behavior with generative AI.
- Enables users to guide LLM-generated text more precisely, providing clear directions for the future design of collaborative writing tools.
-
Experimental or Evaluation Results:
- The acceptance rate for multiple suggestions (74%) was significantly higher than for single suggestions (55%). The acceptance rate for single suggestions improved to 65% when non-diegetic prompts were allowed.
- The quality of user-generated text was strongly correlated with interaction choices; for instance, users who accepted multiple suggestions typically produced longer and more coherent texts.
-
Limitations and Future Directions:
- Limitations:
- Non-diegetic prompts require additional cognitive effort from users, potentially hindering efficiency.
- The short duration of the experiment limits observations of long-term usage habits and changes in prompt preferences.
- GPT-3 suggestions may exhibit repetition or inconsistency.
- Future Directions:
- Develop smarter user interfaces to enable smoother transitions between diegetic and non-diegetic prompts.
- Explore prompt interaction designs in other generative domains (e.g., visual generation systems).
- Investigate the timing of suggestion triggers and users' underlying writing strategies.
- Limitations:
This work provides valuable insights into the integration of interaction design and language models, uncovering user preferences and behavior patterns through experimentation and pointing the way toward optimizing LLM-assisted writing tools.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- When writing with LLMs, how do users choose between diegetic prompts that engage text content and non-diegetic direct instruction prompts?Category: Writing Collaboration, Summarization, and Text SuggestionsSimilar questionsarrow_forward
- What is the impact of different prompting approaches (single versus multiple suggestions) on user writing experience and generated text quality?Category: Writing Collaboration, Summarization, and Text SuggestionsSimilar questionsarrow_forward
- How do users adjust prompt trigger timing during writing tasks to optimize generation outcomes?Category: Writing Collaboration, Summarization, and Text SuggestionsSimilar questionsarrow_forward
Practical Problems
1- Users struggle to achieve natural, fluent, and efficient collaboration when writing with AI-generated text.Category: Writing Collaboration, Summarization, and Text SuggestionsSimilar questionsarrow_forward
- 67%
PhraseFlow: Designs and Empirical Studies of Phrase-Level Input
CHI '21· Generative AI (Text, Image, Music, Video) +2
- 67%
Where Are We So Far? Understanding Data Storytelling Tools from the Perspective of Human-AI Collaboration
CHI '24· Human-LLM Collaboration +1
- 67%
Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models
CHI '24· Human-LLM Collaboration +1
- 67%
Art or Artifice? Large Language Models and the False Promise of Creativity
CHI '24· Human-LLM Collaboration +1
- 67%
Plume: Scaffolding Text Composition in Dashboards
CHI '25· Human-LLM Collaboration +1
- 60%
Tap&Say: Touch Location-Informed Large Language Model for Multimodal Text Correction on Smartphones
CHI '25· Human-LLM Collaboration
Based on Jaccard similarity of research subtopics & professions (≥60%)