RAVEN: Realtime Accessibility in Virtual ENvironments for Blind and Low-Vision People

Voice AccessibilityImmersion & Presence ResearchGenerative AI (Text, Image, Music, Video)Human-LLM CollaborationExplainable AI (XAI)Physicians, Nurses & CliniciansUI/UX DesignersHCI Researchers

Paper Title

RAVEN: Realtime Accessibility in Virtual ENvironments for Blind and Low-Vision People

Publication Info

  • Topic area: Accessibility in virtual 3D environments for blind and low-vision users.
  • Keywords: Accessibility, virtual environments, blind and low-vision users, generative AI, natural language interaction, runtime modification, semantic scene graph, conversational programming.

Background and Problem

  • Problem / challenge: Existing accessibility tools for virtual 3D environments are static, developer-driven, and lack support for dynamic, user-specific adaptations. They often impose steep learning curves and fail to address the nuanced needs of blind and low-vision (BLV) users.
  • Significance: Ensuring equitable access to 3D environments is critical as these spaces become integral to gaming, education, and social interaction. Dynamic, user-driven solutions can empower BLV users and expand their participation.
  • Motivation and related work: Prior work has explored auditory and haptic feedback, visual enhancements, and static accessibility settings in games. However, these approaches are limited in flexibility and personalization. Advances in generative AI and large language models (LLMs) offer opportunities for conversational, runtime scene modifications, which remain underexplored for accessibility.

Solution

  • Proposed approach: RAVEN, a generative AI-powered system enabling BLV users to query and modify 3D virtual environments in real time using natural language.
  • Novelty:
    1. Integration of LLMs with semantic scene data for runtime accessibility modifications.
    2. Introduction of a self-voicing interface for natural language interaction.
    3. Development of accessibility-augmented semantic scene graphs for contextual grounding.
    4. Implementation of prompt-engineering strategies to mitigate hallucinations and enhance reliability.
  • Procedure and key techniques:
    • Users issue natural language prompts for scene queries or modifications.
    • The system retrieves semantic scene data, constructs prompts with accessibility and error-prevention instructions, and uses GROMIT to generate Unity code for modifications.
    • Changes are executed in real time, with spoken feedback provided to users.
    • Iterative refinements were made based on pilot studies, including removing keyboard shortcuts, adding egocentric spatial descriptions, and mitigating hallucinations.

Results

  • Concrete findings:
    • 75.3% of user prompts were successfully executed, with an average response time of 3.1 seconds.
    • System usability was rated highly (SUS score: 79.7), with participants finding it intuitive (M=4.3/5) and empowering.
    • Developers rated the system’s learnability (M=4.7/5) and usability (M=4.3/5) positively.
  • Advantage over baselines:
    • Unlike static, developer-defined tools, RAVEN enables dynamic, user-driven modifications tailored to individual needs.
    • Supports both querying and runtime modifications, addressing a broader range of accessibility challenges.
  • Experiments / evaluation:
    • User study with 8 BLV participants across three scenarios (guided tutorial, task-driven exploration, open-ended exploration) to evaluate usability and interaction strategies.
    • Preliminary developer study with 6 Unity developers to assess integration effort and scalability.
  • Limitations and future work:
    • Non-trivial error rate (22% failed prompts) and reliance on the same LLM for generation and verification.
    • Limited understanding of object affordances, restricting functional interactions.
    • Small sample size and controlled study settings; real-world deployments in commercial games are needed.
    • Future work includes improving error prevention, automating metadata generation, and supporting richer semantic modeling.

Summary

RAVEN introduces a novel approach to accessibility in virtual 3D environments, enabling BLV users to query and modify scenes in real time using natural language. The system leverages generative AI, semantic scene graphs, and prompt-engineering strategies to provide dynamic, user-driven accessibility. Evaluations with BLV participants and Unity developers demonstrated its usability, flexibility, and potential to empower users, while also highlighting challenges with reliability and scalability. Future advancements in error prevention, metadata automation, and real-world deployment will be key to realizing the full potential of generative accessibility systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222233/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791616
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Voice Accessibility, Immersion & Presence Research, Generative AI (Text, Image, Music, Video), Human-LLM Collaboration
work
Professions
Physicians, Nurses & Clinicians, UI/UX Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers