User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI Companions

Explainable AI (XAI)AI Ethics, Fairness & AccountabilityAlgorithmic Fairness & BiasAI/ML Researchers & EngineersHCI Researchers

Research Background and Problem

  • Identified Problems or Challenges: The authors identified that AI companions created based on large language models (LLMs) may generate biased, discriminatory, and harmful statements during daily interactions. These issues can cause psychological harm, reinforce social stereotypes, and have potentially negative impacts, particularly on marginalized groups. Additionally, the general applicability of existing value alignment frameworks led by technical experts (e.g., principles like "helpful, honest, and harmless") is limited, making it difficult to fully meet users' specific needs in real-world interactions.

  • Significance: AI companions have evolved from simple chatbots into virtual friends, partners, and even family members, forming emotional connections with users. This emotional reliance causes people to experience strong emotional distress and harm when encountering discriminatory content generated by AI. Therefore, in dynamic real-world user-AI interaction scenarios, there is a need to introduce a new value alignment approach to prevent and mitigate these issues.

  • Research Motivation and Related Work: The authors propose the concept of User-Driven Value Alignment, aiming to explore how users can actively identify and correct harmful behaviors of AI systems through interaction. This approach seeks to enhance user agency in aligning AI systems with their values and ethical standards, representing an innovative departure from previous expert-led methods or user algorithm auditing approaches.


Solution

  • Proposed Method or Solution: The authors introduce the concept of "User-Driven Value Alignment," emphasizing users' active role in identifying, questioning, and correcting biased outputs from AI during daily interactions. By employing various strategies, users can redirect AI behavior to reflect personal or community values. These strategies are categorized into technical strategies (e.g., regeneration or rollback), argumentative strategies (e.g., expressing anger or reasoning), and role-based strategies (e.g., modifying AI role settings).

  • Innovative Aspects:

    • User Proactivity: Shifts from traditional user feedback on training data to a real-time, proactive value adjustment process.
    • Adaptation to Real-World Contexts: The alignment process is directly based on users' individualized interaction scenarios in real life.
    • Focus on Individual and Community Needs: Better reflects users' specific expectations, values, and sociocultural contexts compared to generic frameworks.
    • Emphasis on Long-Term Iterative Optimization: Users gradually adjust AI behavior through continuous interaction and feedback, promoting behavioral improvement over time.
  • Implementation Steps and Key Techniques:

    1. Collect 77 user-submitted complaint posts about discriminatory statements made by AI companions from various social media platforms (e.g., Reddit, TikTok).
    2. Recruit 20 experienced users of AI companion applications for semi-structured interviews to investigate their strategies for countering AI bias and their understanding of AI behavior.
    3. Use reflexive thematic analysis to code the interview and post data, revealing categories of perceived discriminatory statements, conceptualizations of AI behavior, and user alignment strategies.

Research Findings

  • Specific Findings:

    1. Identified six common types of discrimination perceived by users, including misogyny, LGBTQ+ bias, appearance bias, ableism, racial discrimination, and socioeconomic bias.
    2. Users conceptualized AI behavior in three forms: Machine, Baby (requiring teaching), and Cosplayer (role-playing).
    3. Summarized seven user alignment strategies, categorized into three higher-level types: technical, argumentative, and role-based.
    4. Users' perceptions of the short-term effects of these strategies indicate a gap between reality and expectations in the alignment process.
  • Advantages Over Existing Solutions: Compared to expert-led approaches, user-driven value alignment emphasizes user participation and autonomy, better adapting to the complex and dynamic needs of individuals and communities in real-world scenarios. Additionally, it establishes a more long-term and effective mechanism for behavioral improvement through iterative interactions.

  • Experimental or Evaluation Results: In short-term evaluations, technical strategies often failed to fully resolve issues, while argumentative strategies, though effective, could impose an emotional burden on users. Role-based strategies were generally suitable for short-term corrections but were insufficient to address deeper bias issues.

  • Limitations and Future Directions:

    1. The study relies on self-reported data, which may be subject to emotional and narrative biases.
    2. The sample size is small, primarily employing qualitative methods, limiting generalizability.
    3. The research focuses solely on bias issues in AI companions, with potential for expansion to other types of AI systems and value alignment domains.
    4. The long-term effectiveness of these strategies has not been tested, warranting more targeted large-scale experiments or longitudinal studies.

Conclusion

This study takes an important step toward exploring the phenomenon of user-driven value alignment. By analyzing users' perceptions of AI companion biases, conceptualizations of behavior, and the use of alignment strategies, the authors propose a more adaptive and user-autonomous path for value alignment, highlighting areas for improvement in existing designs. Future work could further quantify the effectiveness of these strategies and expand the concept's applicability to better support the co-evolution of humans and AI systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188861/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713477
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Explainable AI (XAI), AI Ethics, Fairness & Accountability, Algorithmic Fairness & Bias
work
Professions
AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers