Title of the Paper

EyeSayCorrect: Eye Gaze and Voice Based Hands-free Text Correction for Mobile Devices

Paper Information

  • Research Domain: Human-Computer Interaction, specifically text correction technology for mobile devices based on eye tracking and speech recognition.
  • Keywords: Multimodal interaction, eye tracking, text correction, voice input, mobile devices, Bayesian methods, human-computer interaction, accessibility, speech recognition, user study

Research Background and Problem

  • Problem or Challenge:

    1. Text correction on mobile devices typically relies on manual operations, requiring precise cursor positioning, which is user-unfriendly and prone to errors.
    2. Alternative input modes, such as eye tracking and voice input, have limitations—for instance, eye tracking lacks precision, and voice input is sensitive to environmental noise and privacy concerns.
  • Significance:

    • Text correction is a critical function in daily mobile device usage, widely applied in search, emails, and social networks.
    • Providing an effective hands-free text correction method for users with physical disabilities has significant social importance.
  • Research Motivation:

    • Develop a multimodal text correction method that requires no manual intervention, integrating eye gaze tracking and speech recognition to overcome the shortcomings of existing technologies.
    • Introduce a Bayesian model-based error-tolerant text correction strategy to address the lack of precision in eye tracking.

Solution

  • Key Methods and Solutions:

    • Proposed the EyeSayCorrect system, which combines eye gaze tracking and voice input for multimodal text correction.
    • Users select target words by gazing at them and provide new words or phrases via voice input.
    • The system determines the user's correction intent using text context, speech recognition results, and Bayesian methods.
  • Innovations:

    • Introduced a two-dimensional target selection algorithm based on Bayesian theory, which tolerates noise in eye tracking data and supports the selection of smaller text targets.
    • Assigned higher prior probabilities to misspelled words to optimize target selection efficiency.
    • Integrated algorithms based on language models, word embedding, and edit distance to enable intelligent context-aware corrections.
  • Implementation Steps and Technical Details:

    1. Users gaze at a word on the screen, and an area of interest accumulates attention values; once a threshold is reached, the target is selected.
    2. Users input corrected content via voice.
    3. Based on the prior probability of the target text, eye gaze trajectory, and Bayesian inference, the system determines the selected word.
    4. The Apple Speech framework processes speech recognition and generates appropriate correction suggestions using context.

Research Outcomes

  • Specific Outcomes:

    • Verified the feasibility of EyeSayCorrect, demonstrating its ability to achieve hands-free text correction.
    • User experiments showed that assigning higher prior probabilities to misspelled words during the gaze selection phase significantly reduced task completion time.
  • Experiments and Evaluation:

    • Comparative experiments with and without prior conditions showed that prioritizing misspelled words improved user experience:
      • For small font sizes, task completion time was reduced by approximately 23.79%, and word selection time decreased by 40.35%.
      • For large font sizes, task completion time was reduced by about 9.26%.
    • Compared to traditional touch-based methods, EyeSayCorrect's correction speed was about 65% of the touch method but supported hands-free correction functionality.
    • Subjective evaluations indicated that participants preferred the EyeSayCorrect system with misspelling priors, citing reduced physical and psychological burden.
  • Limitations and Future Directions:

    • Limitations:
      • The system relies on voice input, which may raise privacy concerns in public settings.
      • There is room for improvement in the quality of speech recognition results and correction candidates.
    • Future Directions:
      • Optimize Bayesian prior parameters for misspelled words to achieve optimal performance.
      • Explore supplementary input modes, such as introducing more private interaction methods for public environments.
      • Enhance the system's robustness to different user accents and background noise.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/79954/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3490099.3511103
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
16 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Voice User Interface (VUI) Design
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
7 related papers