Research Background and Issues

  • Problem or Challenge: Current conversational AI primarily responds passively to user input. Even in multi-party dialogue scenarios, it mainly relies on dialogue context to predict the next speaker, neglecting AI's proactivity. Moreover, existing methods exhibit limited capability in self-selecting speaking turns during multi-party dialogues, as this requires understanding intrinsic motivations.
  • Significance: Proactive dialogue enables AI to behave more like humans, contributing meaningful content at appropriate moments without explicit triggers. This is particularly important in multi-party conversations, enhancing the fluidity and naturalness of interactions.
  • Research Motivation and Related Work:
    • Traditional prediction strategies rely on external context (e.g., eye contact or linguistic cues), which often lack naturalness or coherence in scenarios without explicit speaker allocation.
    • Fine-tuned models perform inadequately in open-ended dialogue tasks and frequently fail to surpass simple baseline methods.
    • Inspired by cognitive psychology and linguistics, this study aims to simulate human "internal thinking" (e.g., latent but unexpressed thoughts) to enable AI to proactively engage at suitable moments.

Solution

  • Method or Solution: The "Inner Thoughts" framework is proposed, simulating the process of AI generating "internal thoughts" parallel to ongoing dialogue and deciding participation timing based on intrinsic motivations.
  • Innovations:
    • Introducing an AI intrinsic motivation model to enable AI to transcend mere reactions to external stimuli.
    • Proposing a five-step process from cognitive and linguistic perspectives: "trigger events, long-term memory retrieval, thought generation, motivation evaluation, and participation decision."
    • Accounting for human-like complex factors in multi-party dialogues (e.g., topic relevance, urgency, and information gaps) to drive AI's proactive behavior closer to human-like interaction.
  • Implementation Steps:
    • Trigger: Define events that prompt AI to generate thoughts, such as new messages or pauses in dialogue.
    • Retrieval: Search relevant information from AI's short-term and long-term memory.
    • Thought Generation: Use language models to generate potential internal thoughts.
    • Evaluation: Assess these generated thoughts based on eight intrinsic motivations summarized in the study (e.g., relevance, information gap, impact expectation, urgency).
    • Participation: Decide whether to express a thought or remain silent based on evaluation results, and even interrupt under appropriate motivational conditions.

Research Outcomes

  • Specific Outcomes:
    • Developed the Inner Thoughts Playground system and the Swimmy chatbot for multi-party dialogue simulation.
    • Experimental results demonstrate that this framework significantly outperforms baseline models (based on next-speaker prediction) across multiple metrics, such as turn-taking appropriateness, coherence, human-likeness, proactivity, and adaptability.
  • Comparison with Existing Solutions:
    • Baseline models primarily rely on reactive prediction and often perform poorly in self-selected speaking scenarios. In contrast, the Inner Thoughts framework, through intrinsic motivation modeling, aligns more closely with the natural flow of human dialogue.
  • Experimental or Evaluation Results:
    • In simulated dialogues, the Inner Thoughts framework significantly improved dialogue coherence, team member interaction, and topic expansion.
    • In user experiments, AI with varying proactivity settings was recognized by users, with the "proactive contributor" style being rated as the most natural and preferred.
  • Limitations and Future Directions:
    • Quality control of generated content within inner thoughts remains challenging, as it may produce repetitive or incoherent ideas. Advanced methods like knowledge graphs could be introduced.
    • The current framework relies on fixed proactivity thresholds and does not employ data-driven methods for real-time adjustment.
    • Technical evaluations of the framework are primarily focused on text-based scenarios, necessitating further exploration in real-time audio and multimodal contexts.
    • Automated evaluation methods require further development, as existing approaches are resource-intensive and costly. Future work could explore scalable automated evaluation metrics.

By introducing an intrinsic motivation evaluation mechanism, the Inner Thoughts framework demonstrates profound potential for proactive AI applications, including collaborative tasks, ethical decision-making, and multimodal interactions, establishing a new design benchmark for future proactive conversational AI research.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188461/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713760
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Conversational Chatbots, Agent Personality & Anthropomorphism, Human-LLM Collaboration
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
10 related papers