Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMs
Authors
Title of the Paper
Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt-Based Interactions with LLMs
Paper Information
- Subject Areas: Human-Computer Interaction (HCI), Large Language Models (LLMs), Cognitive Psychology
- Keywords: Large Language Models (LLMs), Prompt Interaction, Cognitive Psychology, Interface Design, Intent Generation, Cognitive Models, Artificial Intelligence
Research Background and Problem Statement
-
Identified Problems or Challenges:
- Users face significant cognitive challenges when interacting with Large Language Models (LLMs), particularly in clearly articulating their goals to achieve high-quality system outputs.
- Current interaction systems and models (e.g., Norman's theory of execution and evaluation gulfs) inadequately address the process of intent formation as users transition from goals to actions.
- Due to the dynamic generative capabilities and high flexibility of LLMs, users lack a clear mental model of how these models work, making it difficult to design effective prompts and leading to inconsistent output quality.
-
Importance of the Problem:
- LLMs demonstrate immense potential in complex tasks such as content generation, writing assistance, and code development. However, this potential is limited if users cannot design effective prompts.
- Poor or low-quality outputs can hinder user efficiency, increase cognitive load, and, in some cases (e.g., medical or military applications), result in errors or harmful consequences.
-
Research Motivation and Related Work:
- The primary motivation is to improve interface design by deeply understanding the connection between cognitive processes and human-computer interaction, enabling users to better leverage LLM capabilities.
- Previous studies have explored prompt engineering, limitations of LLM interactions (e.g., lack of interpretability, misleading outputs), and adapting to user needs. However, few have focused on cognition and intent construction.
Proposed Solution
-
Proposed Methods or Solutions:
- The authors propose a new model—"Intent as Action"—which redefines the interaction paradigm between users and LLMs by treating "intent construction as an action."
- They introduce the "Model for Bridging the Gulf of Envisioning" and identify three specific cognitive gaps in LLM interactions:
- Capability Gap: Users lack an understanding of the tasks LLMs can perform.
- Instruction Gap: Users struggle to effectively express their intent through textual prompts.
- Intentionality Gap: Users have unclear expectations of the desired output, making it difficult to evaluate generated content.
-
Innovative Contributions:
- The concept of the "Gulf of Envisioning" is introduced as an interaction stage not fully addressed in Norman's traditional model.
- A set of specific design patterns (e.g., diverse output options, visualized prompt-generation trajectories) is proposed to help users construct prompts, plan tasks, and evaluate outputs more effectively.
- The study emphasizes the dynamic and cognitively complex nature of LLM interactions, addressing the discrepancies between users' mental models and generative systems.
-
Implementation Steps and Techniques:
- Design a semi-structured interface with structured templates to help users clarify their prompts.
- Support divergent thinking and display multiple output options (e.g., through graph-based nodes) to help users manage different output paths.
- Introduce interpretability mechanisms to make LLM outputs more understandable, such as adding annotations in code-generation scenarios or allowing users to directly edit generated content.
Research Outcomes
-
Specific Outcomes:
- Systematically constructed a new cognitive model for human-LLM interaction, identifying the "Gulf of Envisioning" and its associated three gaps.
- Analyzed user cognitive challenges in various generative tasks using examples like ChatGPT, Spellburst, and Cursor, and proposed strategies to address these challenges.
- Provided six design patterns, including diverse output displays and domain-specific prompting strategies, to help users bridge interaction gaps with LLMs.
-
Comparison with Existing Solutions and Advantages:
- Compared to traditional direct-manipulation interfaces, the proposed patterns are better suited to handling the dynamic functionalities and output uncertainties of LLMs.
- The proposed design solutions focus more on users' cognitive processes, emphasizing support for task planning, prompt generation, and output evaluation.
-
Experimental or Evaluation Results:
- The paper analyzed the limitations of existing systems (e.g., ChatGPT, Spellburst, and Cursor) through case studies and proposed future optimization directions based on these limitations.
-
Limitations and Future Directions:
- The current model is primarily based on theoretical derivation, and its practical application and validation require more user studies to quantify improvements in user experience.
- Different user groups (e.g., novices vs. experts) may have varying understandings and adaptations to the proposed design patterns, necessitating further research to develop targeted guidance strategies.
- Bridging the gap between LLM generative capabilities and users' mental models requires further exploration of more efficient visualization and interpretability tools.
Conclusion
- Contributions: This study delves into the core cognitive challenges in user-LLM interactions, proposing specific strategies and design patterns to bridge cognitive gaps and laying a foundation for future research in interaction design.
- Future Outlook: Further research into strategies for constructing user mental models, improving the transparency of generative model interactions, and evaluating more complex multimodal interaction systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What cognitive challenges do users face when interacting with LLMs, particularly in clearly expressing intent?Category: Human-AI Collaborative Optimization and Preference AlignmentSimilar questionsarrow_forward
- How can interface design and cognitive models help users more efficiently generate effective prompts?Category: Human-AI Collaborative Optimization and Preference AlignmentSimilar questionsarrow_forward
- How does users' intent construction process when using LLMs affect output quality, and how can cognitive gaps be bridged?Category: Human-AI Collaborative Optimization and Preference AlignmentSimilar questionsarrow_forward
Practical Problems
1- Users struggle to design high-quality LLM prompts, leading to inconsistent outputs and low efficiency.Category: Human-AI Collaborative Optimization and Preference AlignmentSimilar questionsarrow_forward
- 80%
'It Is Not Always Discovery Time': Four Pragmatic Approaches in Designing AI Systems
CHI '22· Human-LLM Collaboration +2
- 75%
Cells, Generators, and Lenses: Design Framework for Object-Oriented Interaction with Large Language Models
UIST '23· Human-LLM Collaboration
- 67%
Teaching-Learning Interaction: A New Concept for Interaction Design to Support Reflective User Agency in Intelligent Systems
DIS '21· Human-LLM Collaboration +2
- 60%
Mapping Machine Learning Advances from HCI Research to Reveal Starting Places for Design Innovation
CHI '18· Human-LLM Collaboration
- 60%
Sketching NLP: A Case Study of Exploring the Right Things To Design with Language Intelligence
CHI '19· Human-LLM Collaboration +1
- 60%
An Evaluation of Situational Autonomy for Human-AI Collaboration in a Shared Workspace Setting
CHI '24· Human-LLM Collaboration +1
- 60%
IntentTuner: An Interactive Framework for Integrating Human Intentions in Fine-tuning Text-to-Image Generative Models
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 60%
Deus Ex Machina and Personas from Large Language Models: Investigating the Composition of AI-Generated Persona Descriptions
CHI '24· Human-LLM Collaboration +1
- 60%
User Experience of LLM-based Recommendation Systems: A Case of Music Recommendation
CHI '25· Human-LLM Collaboration +1
- 60%
AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools
CHI '25· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)