Emotional Dialogue Generation Using Image-Grounded Language Models

Intelligent Voice Assistants (Alexa, Siri, etc.)Agent Personality & Anthropomorphism

Computer-based conversational agents are becoming ubiquitous. However, for these systems to be engaging and valuable to the user, they must be able to express emotion, in addition to providing informative responses. Humans rely on much more than language during conversations; visual information is key to providing context. We present the first example of an image-grounded conversational agent using visual sentiment, facial expression and scene features. We show that key qualities of the generated dialogue can be manipulated by the features used for training the agent. We evaluate our model on a large and very challenging real-world dataset of conversations from social media (Twitter). The image-grounding leads to significantly more informative, emotional and specific responses, and the exact qualities can be tuned depending on the image features used. Furthermore, our model improves the objective quality of dialogue responses when evaluated on standard natural language metrics.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/5723/2018

AdRecommended

Learn AI Coding at CodeNow

At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2018
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Agent Personality & Anthropomorphism
work
Professions
—
article
Content Status
Abstract only
hub
Related Papers
10 related papers