Navigates Like Me: Understanding How People Evaluate Human-Like AI in Video Games

Human-LLM CollaborationGame UX & Player BehaviorGamification DesignGame Developers & DesignersEsports AthletesHCI Researchers

Title of the Paper

Navigates Like Me: Understanding How People Evaluate Human-Like AI in Video Games

Paper Information

  • Research Area: Applications of Artificial Intelligence and Human-Computer Interaction (HCI) in Video Games
  • Keywords: Human behavior, Turing Test, AI navigation, game agents, user study, video games, human-computer interaction design, behavior analysis, reinforcement learning, intelligent agents

Research Background and Problem Statement

  • Identified Issues or Challenges:
    • Current non-player character (NPC) behaviors often rely on predefined rules, lacking sufficient "human-like" qualities.
    • The usability and realism of AI in games are constrained by the complexity of behavior design.
    • There is no unified standard for evaluating whether AI exhibits human-like characteristics.
  • Significance:
    • Realistic NPC behaviors are crucial for immersive gaming experiences, enhancing player interaction and engagement.
    • In broader contexts (e.g., autonomous driving, collaborative robots), the "human-like" quality of AI is also a key factor for its effectiveness.
  • Research Motivation and Related Work:
    • Improving the realism of AI behaviors (e.g., navigation) contributes to more natural human-AI collaboration.
    • Previous work proposed a "Human Navigation Turing Test" (HNTT) to evaluate the human-likeness of navigation behaviors, but no AI has fully passed this test.
    • Developing AI navigation agents with more human-like characteristics remains an open challenge.

Proposed Solution

  • Proposed Solution:
    • Develop a reinforcement learning-based "reward-shaping agent" to improve human-like navigation characteristics in video games by adjusting reward signals and expanding the set of behavioral actions.
    • Address specific issues: reduce abrupt camera angle changes, avoid frequent collisions, and enhance the smoothness of movement.
  • Innovations:
    • Employ a simple and intuitive reward-shaping method to effectively guide the AI learning process, enabling smoother and more human-like navigation behaviors.
    • Propose, for the first time, a quantitative statistical standard to test whether an agent passes the HNTT, making the evaluation of human-likeness more scientific.
  • Implementation Steps:
    1. Analyze the behavioral patterns of existing baseline agents (symbolic agent and hybrid agent) to identify their shortcomings.
    2. Design a reward-shaping strategy: penalize undesirable behaviors and reward target behaviors, such as reducing penalties for camera swings, adding collision penalties, and rewarding time-efficient navigation.
    3. Expand the AI's action set to enable more natural turning movements.
    4. Conduct large-scale user studies using Amazon Mechanical Turk to evaluate the similarity between the improved navigation agent and human navigation behaviors.

Research Findings

  • Specific Findings:
    • The newly developed reward-shaping agent was the first to pass the HNTT, being deemed indistinguishable from human navigation behaviors; the two baseline agents failed the test.
    • The AI's behavior was widely rated by human participants as smoother, more goal-oriented, and better at responding to the environment.
    • Key behavioral features influencing the evaluation of "human-likeness" (e.g., smooth motion, collision avoidance, adaptability to the environment) were analyzed.
  • Advantages Compared to Existing Solutions:
    • The proposed reward-shaping agent more efficiently simulates human navigation characteristics compared to baseline agents, with participants unable to reliably distinguish the new agent from human behavior.
    • The evaluation process significantly reduced subjective bias through clear statistical metrics.
  • Experimental and Evaluation Results:
    • Sample sizes: Human vs. Symbolic (50 participants), Human vs. Hybrid (50 participants), Human vs. Reward-shaping (92 participants).
    • Under the "reward-shaping agent" condition, the median detection accuracy of evaluators was 0.50 (95% confidence interval includes 0.5), indicating that the AI behavior could not be significantly distinguished.
    • Analyses of rigidity, smoothness, and goal orientation validated the effectiveness of the reward-shaping strategy.
  • Limitations and Future Directions:
    • The current study focuses on third-person perspective game navigation tasks, and its conclusions may not directly generalize to other complex interactive environments (e.g., driving simulations or open-world games).
    • The proposed navigation characteristic model can serve as a preliminary framework for more complex future studies, including cross-cultural evaluations.
    • Future research should enhance interpretability tools for AI design and support personalized participant needs, further improving interactivity and educational value.

Through the above analysis, this paper provides critical guidance for advancing the deep integration of AI and human behavior, while establishing a reproducible and scalable evaluation and optimization methodology.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96541/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581348
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
10 authors
sell
Subtopics
Human-LLM Collaboration, Game UX & Player Behavior, Gamification Design
work
Professions
Game Developers & Designers, Esports Athletes, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers