Artificial Intimacy: Exploring Normativity and Personalization through Fine-tuning LLM Chatbots
Authors
Research Background and Issues
-
What problems or challenges did the authors identify?
- The content generated by large language models (LLMs) often avoids discussing sensitive topics, addressing controversial prompts with evasive or generic non-answers, potentially leading to bias and erasure of diversity.
- Current LLM-based chatbots are being used in attempts at companion-like emotional interactions but lack the ability for empathy, care, and historical memory, resulting in artificial models that only "imitate humans" emotionally.
- Developers attempt to align LLMs with values through fine-tuning, but this process often yields ambiguous results and is fraught with ethical controversies stemming from design decisions, such as forced value outputs or emotional manipulation without consent.
-
Why is this issue important? As the commercialization of LLM-driven chatbots continues to grow, these technologies are being used to mimic emotional relationships and even claim to alleviate loneliness. However, such artificial emotional connections have sparked widespread concerns about emotional manipulation, inauthenticity, and the potential negative impacts of value alignment methods.
-
Research Motivation and Related Work The authors emphasize the ethical challenges implicit in large language models: since OpenAI launched models like GPT-3, the industry and academia have recognized the toxic biases in these technologies, as well as the erasure of minority voices through content "sanitization" when handling sensitive topics. The paper aims to explore whether "extreme personalization" through fine-tuning can address these issues and investigates the ethical and emotional complexities of interacting with artificial intelligence.
Solutions
-
What methods or solutions did the authors propose?
- Using artistic research methods, fine-tuning the GPT-3 model with users' social media data to create chatbots in an extremely personalized manner.
- Developing an artwork titled "Artificial Intimacy," which combines fine-tuned chatbots with artistic installations to simulate value alignment systems based on non-traditional identities (e.g., LGBTQ and BIPOC).
- Conducting participatory workshops where users generate their own "you-bots," analyzing their performance and interactions with participants to observe how personalization and value embodiment affect user experiences.
-
What is innovative about this solution?
- Exploration of extreme personalization: Extracting personal values from users' social media data and fine-tuning individual data to study whether artificial intelligence mimicking self-identity could provoke artificial emotional intimacy.
- Integration of art and technology: The authors use art as a framework for reflection, enabling the ethical consequences of fine-tuning technology to be showcased and discussed through artistic works.
- Diversity in dialogue: Injecting non-mainstream values (e.g., race, gender, sexual orientation) into chatbots to explore feasibility and ethical challenges.
-
What are the implementation steps and key technologies used?
- Utilizing GPT-3's fine-tuning capabilities, importing user-provided social media data in JSONL format to generate personalized language models.
- Creating two chatbots based on non-mainstream identities, followed by exhibitions and public interactions.
- Organizing four workshops where participants interact with both the general GPT-3 model and fine-tuned "you-bots" and "Leslie-bot," comparing the behavioral characteristics of different models.
Research Outcomes
-
What specific outcomes were achieved?
- When interacting with the general GPT-3 model, participants found its responses politically correct and neutral but lacking depth and personalization; its avoidance of controversial topics weakened emotional interaction between users and the model.
- Fine-tuned models demonstrated stronger personalization and subjectivity, making it easier for users to "personify" the model, but also elicited discomfort, boredom, and humor, especially in cases of value conflicts.
- The study revealed that fine-tuning based on user data led to the formation of artificial intimacy, with users developing unrealistic emotional expectations of the model, such as expecting it to exhibit care or empathy.
- The unpredictability of the model's outputs (e.g., humorous expressions or emotional mimicry) enhanced users' sense of identification, but the "self-mirroring" characteristic of mimicry was deemed "creepy" by some users.
-
What advantages does it have compared to existing solutions?
- Compared to general models, fine-tuned models are more vivid and personalized in interactions, significantly improving diversity and simulated empathy (though artificial).
- The research design, which integrates artistic methods, provides deeper analysis of ethical and social issues, offering a novel perspective on human-machine relationships.
-
What were the experimental or evaluation results?
- Value alignment in interactions: Fine-tuned models reflected the personality or worldview of the data provider, with "you-bots" strongly mimicking the themes and tone of users' social media data.
- Users were more inclined to assign personal gender or anthropomorphic traits to fine-tuned models, with this "personification" helping them understand interaction content but potentially leading to risks of trust and emotional dependency.
- Value conflict issues: For example, conversations with the Leslie-bot, based on unconventional values, often triggered discomfort and rejection from users. The authors argue that such uncompromising value imposition is not conducive to fostering deep reflection.
-
Limitations and Future Directions
- Limitations: The fine-tuning of models relies on the quality of user-provided data, and due to GPT-3's scale limitations, fine-tuned models still face issues with semantic coherence and content accuracy.
- Future Directions: The study highlights the need for more refined and dynamic methods to capture human values and suggests designing more participatory and contextualized model training approaches for value alignment, while further exploring ways to mitigate ethical risks of fine-tuned models (e.g., emotional manipulation).
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Can extreme personalization through fine-tuning mitigate bias and lack of diversity in LLMs on sensitive topics?Category: Bias and Fairness in Large Language Models and GenAISimilar questionsarrow_forward
- Can robots fine-tuned on users' social media data elicit perceived emotional intimacy?Category: Bias and Fairness in Large Language Models and GenAISimilar questionsarrow_forward
- How are non-mainstream values embodied through fine-tuned models, and what impact do they have on users?Category: Bias and Fairness in Large Language Models and GenAISimilar questionsarrow_forward
Practical Problems
1- Users often find emotional interaction with LLM-based chatbots cold and inauthentic.Category: Bias and Fairness in Large Language Models and GenAISimilar questionsarrow_forward
- 100%
Will AI Console Me when I Lose my Pet? Understanding Perceptions of AI-Mediated Email Writing
CHI '22· Agent Personality & Anthropomorphism +2
- 67%
Owning Mistakes Sincerely: Strategies for Mitigating AI Errors
CHI '22· Agent Personality & Anthropomorphism +1
- 67%
Do You Mind? User Perceptions of Machine Consciousness
CHI '23· Agent Personality & Anthropomorphism +1
- 67%
CloChat: Understanding How People Customize, Interact, and Experience Personas in Large Language Models
CHI '24· Agent Personality & Anthropomorphism +1
- 67%
Relational AI: Facilitating Intergroup Cooperation with Socially Aware Conversational Support
CHI '25· Agent Personality & Anthropomorphism +1
- 67%
Beyond the Illusion: LLMs and the Case for Pragmatic Cues in Conversation
CUI '25· Agent Personality & Anthropomorphism +1
- 67%
Can Large Language Models Be Good Companions? An LLM-Based Eyewear System with Conversational Common Ground
UbiComp '24· Agent Personality & Anthropomorphism +1
Based on Jaccard similarity of research subtopics & professions (≥60%)