The ORBIT India Dataset: Understanding the Challenges of Collecting a Disability-First AI Dataset in Low-Resource Environments

Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Cognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia)Generative AI (Text, Image, Music, Video)Low-Resource Languages & Digital InclusionPhysicians, Nurses & CliniciansAssistive Technology SpecialistsHCI Researchers

Paper Title

The ORBIT India Dataset: Understanding the Challenges of Collecting a Disability-First AI Dataset in Low-Resource Environments

Publication Info

  • Topic area: AI dataset collection for disability-first object recognition in low-resource settings.
  • Keywords: Disability-first AI, teachable object recognition, low-resource environments, ORBIT-India dataset, Global South, dataset collection, cultural adaptation, blind and low vision, PII detection, inclusive AI.

Background and Problem

  • Problem / challenge: Existing AI systems for object recognition are predominantly trained on datasets from the Global North, leading to poor performance in culturally diverse and low-resource contexts like India. Additionally, datasets often fail to reflect the lived realities of blind and low-vision users, whose images include challenges like low lighting, motion blur, and occlusion.
  • Significance: India has the world’s largest blind population, and creating datasets that reflect their environments and needs is critical for equitable AI development. This can improve accessibility and inclusivity in AI systems.
  • Motivation and related work: Prior datasets like ORBIT and VizWiz have made strides in disability-first data collection but are limited to high-resource contexts. Efforts to adapt such protocols to the Global South remain underexplored, necessitating a focus on cultural, technological, and socio-economic factors unique to these regions.

Solution

  • Proposed approach: The ORBIT-India dataset, a teachable object recognition dataset collected by blind and low-vision individuals in India, using an adapted version of the Find My Things app.
  • Novelty:
    1. Introduction of the ORBIT-India dataset with 105,243 images of 76 objects, reflecting Indian material culture and collected by 12 blind or low-vision individuals.
    2. Adaptation of a Global North dataset collection protocol to the Indian context, addressing challenges like device variability, connectivity, and cultural differences.
    3. Recommendations for inclusive, cross-geography dataset collection practices, emphasizing privacy, accessibility, and cultural relevance.
  • Procedure and key techniques:
    • Data collectors used the Find My Things app to record videos of personal objects, with modifications for Android compatibility, Hindi localization, and culturally relevant instructions.
    • Videos were categorized into clean (training) and cluttered (testing) scenarios, with privacy checks and annotations for personally identifiable information (PII).
    • Feedback and support mechanisms, including “Datathons” and iterative reviews, ensured high-quality data collection.

Results

  • Concrete findings:
    • Dataset includes 105,243 images from 587 videos, with 75,994 clean images and 29,249 cluttered images.
    • 3.8% of frames contained PII, and 1.11% had the object completely out of frame.
    • Cultural uniqueness: Objects like bangles, hair oil, and steel glasses reflect Indian material culture, while naming conventions and household layouts highlight regional diversity.
  • Advantage over baselines:
    • Unlike prior datasets, ORBIT-India captures the lived realities of Indian blind and low-vision users, including culturally specific objects and environmental conditions.
    • Provides a foundation for adapting AI systems to low-resource and culturally diverse contexts.
  • Experiments / evaluation:
    • Data collection involved 12 participants (average age 27) from urban India, contributing an average of 8,770 images each.
    • Feedback interviews revealed high usability of the app and valuable insights into object selection, filming, and privacy considerations.
  • Limitations and future work:
    • Dataset size is smaller compared to other visual datasets, and participants were predominantly urban, young, and male, limiting generalizability.
    • Future work should expand to rural areas, diverse socio-economic groups, and larger cohorts, while addressing device and connectivity constraints.

Summary

The ORBIT-India dataset is the first teachable object recognition dataset collected by blind and low-vision individuals in India, comprising 105,243 images of 76 objects. It reflects Indian material culture and addresses challenges in adapting dataset collection protocols from the Global North to low-resource settings. The study highlights the importance of culturally informed, inclusive practices in AI dataset collection and provides recommendations for future efforts. While the dataset is a meaningful first step, future work should focus on scaling and diversifying participation to better represent India’s socio-cultural diversity.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222651/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791099
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille), Cognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia), Generative AI (Text, Image, Music, Video), Low-Resource Languages & Digital Inclusion
work
Professions
Physicians, Nurses & Clinicians, Assistive Technology Specialists, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers