User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI Companions
Authors
Research Background and Problem
-
Identified Problems or Challenges: The authors identified that AI companions created based on large language models (LLMs) may generate biased, discriminatory, and harmful statements during daily interactions. These issues can cause psychological harm, reinforce social stereotypes, and have potentially negative impacts, particularly on marginalized groups. Additionally, the general applicability of existing value alignment frameworks led by technical experts (e.g., principles like "helpful, honest, and harmless") is limited, making it difficult to fully meet users' specific needs in real-world interactions.
-
Significance: AI companions have evolved from simple chatbots into virtual friends, partners, and even family members, forming emotional connections with users. This emotional reliance causes people to experience strong emotional distress and harm when encountering discriminatory content generated by AI. Therefore, in dynamic real-world user-AI interaction scenarios, there is a need to introduce a new value alignment approach to prevent and mitigate these issues.
-
Research Motivation and Related Work: The authors propose the concept of User-Driven Value Alignment, aiming to explore how users can actively identify and correct harmful behaviors of AI systems through interaction. This approach seeks to enhance user agency in aligning AI systems with their values and ethical standards, representing an innovative departure from previous expert-led methods or user algorithm auditing approaches.
Solution
-
Proposed Method or Solution: The authors introduce the concept of "User-Driven Value Alignment," emphasizing users' active role in identifying, questioning, and correcting biased outputs from AI during daily interactions. By employing various strategies, users can redirect AI behavior to reflect personal or community values. These strategies are categorized into technical strategies (e.g., regeneration or rollback), argumentative strategies (e.g., expressing anger or reasoning), and role-based strategies (e.g., modifying AI role settings).
-
Innovative Aspects:
- User Proactivity: Shifts from traditional user feedback on training data to a real-time, proactive value adjustment process.
- Adaptation to Real-World Contexts: The alignment process is directly based on users' individualized interaction scenarios in real life.
- Focus on Individual and Community Needs: Better reflects users' specific expectations, values, and sociocultural contexts compared to generic frameworks.
- Emphasis on Long-Term Iterative Optimization: Users gradually adjust AI behavior through continuous interaction and feedback, promoting behavioral improvement over time.
-
Implementation Steps and Key Techniques:
- Collect 77 user-submitted complaint posts about discriminatory statements made by AI companions from various social media platforms (e.g., Reddit, TikTok).
- Recruit 20 experienced users of AI companion applications for semi-structured interviews to investigate their strategies for countering AI bias and their understanding of AI behavior.
- Use reflexive thematic analysis to code the interview and post data, revealing categories of perceived discriminatory statements, conceptualizations of AI behavior, and user alignment strategies.
Research Findings
-
Specific Findings:
- Identified six common types of discrimination perceived by users, including misogyny, LGBTQ+ bias, appearance bias, ableism, racial discrimination, and socioeconomic bias.
- Users conceptualized AI behavior in three forms: Machine, Baby (requiring teaching), and Cosplayer (role-playing).
- Summarized seven user alignment strategies, categorized into three higher-level types: technical, argumentative, and role-based.
- Users' perceptions of the short-term effects of these strategies indicate a gap between reality and expectations in the alignment process.
-
Advantages Over Existing Solutions: Compared to expert-led approaches, user-driven value alignment emphasizes user participation and autonomy, better adapting to the complex and dynamic needs of individuals and communities in real-world scenarios. Additionally, it establishes a more long-term and effective mechanism for behavioral improvement through iterative interactions.
-
Experimental or Evaluation Results: In short-term evaluations, technical strategies often failed to fully resolve issues, while argumentative strategies, though effective, could impose an emotional burden on users. Role-based strategies were generally suitable for short-term corrections but were insufficient to address deeper bias issues.
-
Limitations and Future Directions:
- The study relies on self-reported data, which may be subject to emotional and narrative biases.
- The sample size is small, primarily employing qualitative methods, limiting generalizability.
- The research focuses solely on bias issues in AI companions, with potential for expansion to other types of AI systems and value alignment domains.
- The long-term effectiveness of these strategies has not been tested, warranting more targeted large-scale experiments or longitudinal studies.
Conclusion
This study takes an important step toward exploring the phenomenon of user-driven value alignment. By analyzing users' perceptions of AI companion biases, conceptualizations of behavior, and the use of alignment strategies, the authors propose a more adaptive and user-autonomous path for value alignment, highlighting areas for improvement in existing designs. Future work could further quantify the effectiveness of these strategies and expand the concept's applicability to better support the co-evolution of humans and AI systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What types of perceived bias exist in AI companions?Category: Algorithmic Stigmatization and Social HarmSimilar questionsarrow_forward
- How do users actively adjust AI companion values through interaction to reflect personal or community values?Category: Algorithmic Stigmatization and Social HarmSimilar questionsarrow_forward
- What are the short-term effects and limitations of user-driven value adjustment strategies in real-world scenarios?Category: Algorithmic Stigmatization and Social HarmSimilar questionsarrow_forward
Practical Problems
1- Bias and harmful content in AI companions may cause psychological harm and social impact for users.Category: Algorithmic Stigmatization and Social HarmSimilar questionsarrow_forward
- 100%
Beyond Expertise and Roles: A Framework to Characterize the Stakeholders of Interpretable Machine Learning and their Needs
CHI '21· Explainable AI (XAI) +2
- 83%
Explaining Models: An Empirical Study of How Explanations Impact Fairness Judgment
IUI '19· Explainable AI (XAI) +2
- 83%
The Impact of Explanations on Fairness in Human-AI Decision-Making: Protected vs Proxy Features
IUI '24· Explainable AI (XAI) +2
- 80%
"Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI
CHI '21· Explainable AI (XAI) +1
- 80%
Towards Fairness in Practice: A Practitioner-Oriented Rubric for Evaluating Fair ML Toolkits
CHI '21· AI Ethics, Fairness & Accountability +1
- 80%
Jury Learning: Integrating Dissenting Voices into Machine Learning Models
CHI '22· AI Ethics, Fairness & Accountability +1
- 80%
Capable but Amoral? Comparing AI and Human Expert Collaboration in Ethical Decision Making
CHI '22· AI Ethics, Fairness & Accountability +1
- 80%
Fairness Evaluation in Text Classification: Machine Learning Practitioner Perspectives of Individual and Group Fairness
CHI '23· Explainable AI (XAI) +1
- 80%
Out of Context: Investigating the Bias and Fairness Concerns of "Artificial Intelligence as a Service"
CHI '23· AI Ethics, Fairness & Accountability +1
- 80%
“It is currently hodgepodge”: Examining AI/ML Practitioners’ Challenges during Co-production of Responsible AI Values
CHI '23· AI Ethics, Fairness & Accountability +1
Based on Jaccard similarity of research subtopics & professions (≥60%)