Streaming, Fast and Slow: Cognitive Load-Aware Streaming for Efficient LLM Serving
Generative conversational interfaces powered by large language models (LLMs) typically stream output token-by-token at a rate determined by computational budget, often neglecting actual human reading speeds and the cognitive load associated with the content. This mismatch frequently leads to inefficient use of computational resources. For example, in cloud-based services, streaming content faster than users can read appears unnecessary, resulting in wasted computational resources and potential delays for other users, particularly during peak usage periods. To address this issue, we propose an adaptive streaming method that dynamically adjusts the pacing of LLM streaming output in real-time based on inferred cognitive load. Our approach estimates the cognitive load associated with streaming content and strategically slows down the stream during complex or information-rich segments, thereby freeing computational resources for other users. We conducted a statistical analysis and simulation based on a statistical model derived from data collected in a crowdsourced user study across various types of LLM-generated content. Our results show that this adaptive method can effectively reduce computational consumption while largely maintaining streaming speed above user's normal reading speed.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 60%
Discovering the Syntax and Strategies of Natural Language Programming with Generative Language Models
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 60%
"What It Wants Me To Say": Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models
CHI '23· Generative AI (Text, Image, Music, Video) +1
- 60%
Human-LLM Collaborative Annotation Through Effective Verification of LLM Labels
CHI '24· Human-LLM Collaboration +1
- 60%
IntentTuner: An Interactive Framework for Integrating Human Intentions in Fine-tuning Text-to-Image Generative Models
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 60%
CreAItive Collaboration? Users' Misjudgment of AI-Creativity Affects Their Collaborative Performance
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 60%
AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 60%
D-Twins: Your Digital Twin Designed for Real-Time Boredom Intervention
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 60%
Finding the Conversation: A Method for Scoring Documents for Natural Conversation Content
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 60%
Fluid Transformers and Creative Analogies: Exploring Large Language Models' Capacity for Augmenting Cross-Domain Analogical Creativity
C&C '23· Generative AI (Text, Image, Music, Video) +1
- 60%
Take It, Leave It, or Fix It: Measuring Productivity and Trust in Human-AI Collaboration
IUI '24· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)