Watch It, Don't Imagine It: Creating a Better Caption-Occlusion Metric by Collecting More Ecologically Valid Judgments from DHH Viewers

Voice AccessibilityDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Universal & Inclusive DesignSpeech-Language Pathologists & AudiologistsSpecial Education TeachersDisability Service Providers

Title of the Paper

Watch It, Don’t Imagine It: Creating a Better Caption-Occlusion Metric by Collecting More Ecologically Valid Judgments from DHH Viewers

Paper Information

  • Field of Study: Accessibility Design, Human-Computer Interaction
  • Keywords: Accessibility, Caption Occlusion, Metrics, Regression, Deaf and Hard of Hearing (DHH)

Research Background and Problems

  • What problems or challenges did the authors identify?

    • Television captions may obscure other on-screen content (e.g., text, facial expressions), reducing the viewing experience for Deaf and Hard of Hearing (DHH) individuals.
    • Existing caption quality assessment metrics primarily focus on caption text accuracy, neglecting issues of caption placement and occlusion.
    • Previous studies used static diagrams to have participants "imagine" the impact of caption occlusion, a data collection method that may lack ecological validity.
  • Why is this issue important?

    • Over 360 million people worldwide face hearing challenges, and 15% of U.S. adults rely on television captions to access content.
    • Caption occlusion can impact the comprehension and satisfaction of DHH audiences with television programs, thereby lowering the quality of accessible media services.
  • Motivation and Related Work

    • The authors aim to develop a new metric for assessing caption occlusion severity by collecting data in more realistic contexts, guiding automated caption placement and improving the user experience for DHH audiences.

Solutions

  • What methods or solutions did the authors propose?

    • Developed a "Holistic Judgment Model" using regression techniques to summarize subjective ratings from DHH participants after watching dynamic videos.
    • Compared to the previous "Component Judgment Model," employed a more ecologically valid data collection method by having participants watch actual videos rather than imagining occlusion effects based on static diagrams.
  • What is innovative about this solution?

    • Improved data collection by directly involving participants in watching videos and providing holistic quality ratings of captions, avoiding biases from isolated judgments of individual screen areas.
    • Used regression modeling to calculate the impact of caption occlusion on specific screen regions (e.g., eyes, mouths) on overall viewing quality and optimized the weighting of these factors.
  • What are the implementation steps and key technologies used?

    1. Video Stimuli Generation:
      • Selected 104 video clips from six common television genres, including news, weather, sports, emergency announcements, etc.
      • Created four different caption placement versions for each video.
    2. Video Annotation:
      • Annotated the percentage of time and area that each information region was occluded.
    3. Data Collection:
      • Organized 24 DHH participants to remotely watch the videos and provide ratings.
      • Participants answered the subjective question, "How satisfied are you with the caption placement in the video?" on a 10-point scale.
    4. Regression Analysis:
      • Generated an optimal regression model for each television genre.
      • Used feature engineering to select the best occlusion features for each information region.

Research Findings

  • What specific results were achieved?

    • The new model significantly predicted DHH viewers' subjective ratings of caption occlusion quality, performing particularly well in genres like weather news and sports.
    • Compared to previous models, the new model provided more reasonable explanations for the weights of screen region features, reflecting the dynamic changes in real viewing experiences.
  • What advantages does it have over existing solutions?

    • The new ecologically valid data collection approach better reflects real-world viewing scenarios, effectively reducing biases from imagined impacts.
    • The new model improved the accuracy of caption quality assessments by optimizing the weights through regression analysis.
  • What were the experimental or evaluation results?

    • Compared to the previous "Component Judgment Model," the new "Holistic Judgment Model" achieved higher predictive performance, especially for information-dense videos.
    • The new model explained a significant portion of the variance in ratings, such as an adjusted R² of 0.282 for weather news.
  • Limitations and Future Directions

    • Limitations:
      • The study only used short video clips (30 seconds); longer television programs need to be tested in the future.
      • The sample size was small and covered only a limited age range of the DHH population.
      • The study focused on six television genres, which may not represent the needs of other types of content.
      • Other factors affecting caption quality (e.g., synchronization, style) were not considered.
    • Future Directions:
      • Collect larger-scale participant data, covering a broader range of ages and backgrounds.
      • Develop automated video region detection technologies to improve the efficiency of caption placement tools.
      • Expand the study to include other types of television programs to enhance the model's applicability.

Conclusion

This study improved data collection methods by creating a new "Holistic Judgment Model" based on subjective ratings from watching actual videos, significantly enhancing the quality of caption occlusion severity metrics. This work not only provides guidance for optimizing automated caption placement but also contributes theoretical support to research on audience-video interaction, advancing the field of accessibility design.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/71914/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3517681
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Voice Accessibility, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration), Universal & Inclusive Design
work
Professions
Speech-Language Pathologists & Audiologists, Special Education Teachers, Disability Service Providers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers