Disability-First Design and Creation of A Dataset Showing Private Visual Information Collected With People Who Are Blind
Authors
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Cognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia)Universal & Inclusive DesignPrivacy by Design & User ControlPublic Transit OperatorsAI/ML Researchers & EngineersAssistive Technology Specialists
Title of the Paper
Disability-First Design and Creation of A Dataset Showing Private Visual Information Collected with People Who Are Blind
Paper Information
- Subject Area: Human-Computer Interaction, Data Ethics, Privacy Protection, and AI Dataset Design
- Keywords: dataset, accessible design, privacy, private visual data, image description, computer vision, visual impairment, visual assistant, sample learning
Research Background and Problem
- Identified Issues or Challenges:
- Traditional visual datasets often remove or obscure private content to protect privacy, which limits the development of privacy-preserving AI models.
- Common characteristics in visual data used by blind and low-vision individuals (e.g., blurry images, irregular shooting angles) are not adequately represented in existing datasets. This disparity reduces the representation of blind individuals in AI models.
- Blind individuals cannot easily verify whether private content is unintentionally included in images, posing a risk of privacy breaches when participating in visual dataset creation.
- Significance:
- The goal of creating this dataset is to support the development of privacy-preserving technologies, helping approximately 39 million blind individuals worldwide (including about 8 million in the U.S.) independently identify and manage private content in images and videos.
- Inclusive dataset collection and design can advance AI systems in terms of accessibility and privacy protection.
- Research Motivation and Related Work:
- Recent literature emphasizes fairness and inclusivity in AI, with limited research focusing on people with disabilities.
- Existing datasets (e.g., VizWiz) have improved representation of blind individuals' data but deliberately avoid private content, making them unsuitable for training privacy-preserving models.
Proposed Solution
- Proposed Solution:
- Create a "Disability-First" dataset (BIV-Priv) based on the shooting behavior of blind users, incorporating virtual private props such as credit cards, medicine bottles, and bills.
- This approach captures private content using props instead of participants' actual private items, minimizing privacy risks.
- Innovations:
- Data collection using virtual props to reduce the impact on participants' privacy.
- A multidisciplinary perspective combining accessible design, privacy protection, and computer vision to provide diverse private content and realistic shooting conditions for AI training.
- Inclusion of natural shooting characteristics of blind users in the dataset (e.g., poor lighting, objects misaligned in the frame).
- Implementation Steps and Key Technologies:
- Prop Design and Creation: Design 14 types of common private objects as virtual replicas, using generated virtual information (e.g., randomly generated addresses, phone numbers) to simulate real private data.
- Technical Infrastructure: Develop an iOS mobile application supporting voice guidance and real-time text chat for data collection.
- User Research and Data Collection: Conduct preliminary training with 26 blind participants, who completed data collection tasks by capturing 728 images and 728 videos.
- Ethics and Privacy Controls: Ensure the dataset only includes virtual props and remove any inadvertently captured private content from participants.
Research Outcomes
- Specific Outcomes:
- Created the BIV-Priv dataset, comprising 728 images and 728 videos, covering 14 categories of private props.
- The dataset provides diverse privacy-related scenarios for AI model development and reflects the natural shooting behavior of blind users.
- Advantages Compared to Existing Solutions:
- Offers enhanced privacy protection by avoiding the disclosure of participants' real information.
- Incorporates diverse and multidisciplinary design, increasing the dataset's broad applicability.
- Improves responsiveness to the actual needs and privacy concerns of blind users.
- Experimental and Evaluation Results:
- Data contributed by 26 participants, naturally captured by blind users, exhibited distinctive characteristics such as blurriness, irregular lighting, and object misalignment, enriching AI models with realistic scenarios.
- Participants generally reported feeling more comfortable sharing data using virtual props instead of real content and expressed a desire to contribute to the development of more effective privacy-preserving technologies.
- Limitations and Future Directions:
- The data is limited to blind users in the U.S., lacking broader international and diverse population representation.
- The current dataset includes a limited range of content, covering only 14 types of private props.
- Future work could explore data creation methods for other disability groups and expand content categories (e.g., passports, social security numbers). Research should also investigate automated intervention technologies to reduce users' task burden.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can a dataset prioritizing disabled communities' needs be designed to improve development of privacy-preserving AI models?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- What unique features do images and videos generated by BLV users have for privacy-related AI modeling?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- How can privacy protection and data authenticity needs be balanced during data collection?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
lightbulb
Practical Problems
1- BLV users worry about accidentally leaking private content when sharing data.Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3580922
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille), Cognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia), Universal & Inclusive Design, Privacy by Design & User Control
work
Professions
Public Transit Operators, AI/ML Researchers & Engineers, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers