Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMs

Human-LLM CollaborationUser Research Methods (Interviews, Surveys, Observation)UI/UX DesignersAI/ML Researchers & Engineers

Title of the Paper

Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt-Based Interactions with LLMs

Paper Information

  • Subject Areas: Human-Computer Interaction (HCI), Large Language Models (LLMs), Cognitive Psychology
  • Keywords: Large Language Models (LLMs), Prompt Interaction, Cognitive Psychology, Interface Design, Intent Generation, Cognitive Models, Artificial Intelligence

Research Background and Problem Statement

  • Identified Problems or Challenges:

    1. Users face significant cognitive challenges when interacting with Large Language Models (LLMs), particularly in clearly articulating their goals to achieve high-quality system outputs.
    2. Current interaction systems and models (e.g., Norman's theory of execution and evaluation gulfs) inadequately address the process of intent formation as users transition from goals to actions.
    3. Due to the dynamic generative capabilities and high flexibility of LLMs, users lack a clear mental model of how these models work, making it difficult to design effective prompts and leading to inconsistent output quality.
  • Importance of the Problem:

    • LLMs demonstrate immense potential in complex tasks such as content generation, writing assistance, and code development. However, this potential is limited if users cannot design effective prompts.
    • Poor or low-quality outputs can hinder user efficiency, increase cognitive load, and, in some cases (e.g., medical or military applications), result in errors or harmful consequences.
  • Research Motivation and Related Work:

    • The primary motivation is to improve interface design by deeply understanding the connection between cognitive processes and human-computer interaction, enabling users to better leverage LLM capabilities.
    • Previous studies have explored prompt engineering, limitations of LLM interactions (e.g., lack of interpretability, misleading outputs), and adapting to user needs. However, few have focused on cognition and intent construction.

Proposed Solution

  • Proposed Methods or Solutions:

    1. The authors propose a new model—"Intent as Action"—which redefines the interaction paradigm between users and LLMs by treating "intent construction as an action."
    2. They introduce the "Model for Bridging the Gulf of Envisioning" and identify three specific cognitive gaps in LLM interactions:
      • Capability Gap: Users lack an understanding of the tasks LLMs can perform.
      • Instruction Gap: Users struggle to effectively express their intent through textual prompts.
      • Intentionality Gap: Users have unclear expectations of the desired output, making it difficult to evaluate generated content.
  • Innovative Contributions:

    1. The concept of the "Gulf of Envisioning" is introduced as an interaction stage not fully addressed in Norman's traditional model.
    2. A set of specific design patterns (e.g., diverse output options, visualized prompt-generation trajectories) is proposed to help users construct prompts, plan tasks, and evaluate outputs more effectively.
    3. The study emphasizes the dynamic and cognitively complex nature of LLM interactions, addressing the discrepancies between users' mental models and generative systems.
  • Implementation Steps and Techniques:

    1. Design a semi-structured interface with structured templates to help users clarify their prompts.
    2. Support divergent thinking and display multiple output options (e.g., through graph-based nodes) to help users manage different output paths.
    3. Introduce interpretability mechanisms to make LLM outputs more understandable, such as adding annotations in code-generation scenarios or allowing users to directly edit generated content.

Research Outcomes

  • Specific Outcomes:

    1. Systematically constructed a new cognitive model for human-LLM interaction, identifying the "Gulf of Envisioning" and its associated three gaps.
    2. Analyzed user cognitive challenges in various generative tasks using examples like ChatGPT, Spellburst, and Cursor, and proposed strategies to address these challenges.
    3. Provided six design patterns, including diverse output displays and domain-specific prompting strategies, to help users bridge interaction gaps with LLMs.
  • Comparison with Existing Solutions and Advantages:

    • Compared to traditional direct-manipulation interfaces, the proposed patterns are better suited to handling the dynamic functionalities and output uncertainties of LLMs.
    • The proposed design solutions focus more on users' cognitive processes, emphasizing support for task planning, prompt generation, and output evaluation.
  • Experimental or Evaluation Results:

    • The paper analyzed the limitations of existing systems (e.g., ChatGPT, Spellburst, and Cursor) through case studies and proposed future optimization directions based on these limitations.
  • Limitations and Future Directions:

    1. The current model is primarily based on theoretical derivation, and its practical application and validation require more user studies to quantify improvements in user experience.
    2. Different user groups (e.g., novices vs. experts) may have varying understandings and adaptations to the proposed design patterns, necessitating further research to develop targeted guidance strategies.
    3. Bridging the gap between LLM generative capabilities and users' mental models requires further exploration of more efficient visualization and interpretability tools.

Conclusion

  • Contributions: This study delves into the core cognitive challenges in user-LLM interactions, proposing specific strategies and design patterns to bridge cognitive gaps and laying a foundation for future research in interaction design.
  • Future Outlook: Further research into strategies for constructing user mental models, improving the transparency of generative model interactions, and evaluating more complex multimodal interaction systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147633/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642754
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, User Research Methods (Interviews, Surveys, Observation)
work
Professions
UI/UX Designers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers