Desirable Unfamiliarity: Insights from Eye Movements on Engagement and Readability of Dictation Interfaces

Language Model-Assisted Text InputAI-Assisted Writing & Text GenerationEye Tracking & Gaze InteractionHCI ResearchersUniversity Professors & ResearchersSoftware Engineers & Developers

Paper Title

Desirable Unfamiliarity: Insights from Eye Movements on Engagement and Readability of Dictation Interfaces

Publication Info

  • Topic area: Eye-tracking analysis of dictation interfaces for speech-to-text systems.
  • Keywords: Dictation interfaces, speech-to-text, eye-tracking, reading behavior, LLM-generated summaries, keyword highlighting, user engagement, cognitive load, text readability, transcription systems.

Background and Problem

  • Problem / challenge: Raw speech-to-text (STT) transcripts are often hard to read due to disfluencies, recognition errors, and verbosity. While LLM-based corrections can improve readability, they risk altering the original meaning or causing distractions during speech production.
  • Significance: Improving the readability and usability of dictation interfaces is critical for their adoption in personal, professional, and educational contexts.
  • Motivation and related work: Existing research has focused on transcription accuracy and disfluency correction but has not evaluated the effectiveness of readability-enhancing solutions like keyword highlighting or LLM-generated summaries. Eye-tracking studies on reading behavior in dictation interfaces remain unexplored.

Solution

  • Proposed approach: Conduct an eye-tracking study to evaluate five dictation interfaces: PLAIN, AOC, RAKE, GP-TSM, and SUMMARY, each differing in fidelity to original speech, review aids, and interface responsiveness.
  • Novelty:
    1. Introduced the concept of "Desirable Unfamiliarity," where LLM-generated summaries improve readability despite introducing unfamiliar words.
    2. Compared extractive (RAKE, GP-TSM) and abstractive (SUMMARY) methods for aiding text review in dictation interfaces.
    3. Analyzed eye-tracking data to understand visual engagement and reading effort during speech production and review.
    4. Provided actionable design implications for future dictation interfaces.
  • Procedure and key techniques:
    • Participants (N=20) composed and reviewed social media posts using five interfaces while eye movements were tracked.
    • Interfaces included real-time transcription (PLAIN), periodic corrections (AOC), keyword highlighting (RAKE), grammar-preserving highlights (GP-TSM), and LLM-generated summaries (SUMMARY).
    • Measured gaze engagement, regressions, and fixation metrics during production and review phases.
    • Conducted qualitative analysis through surveys and interviews.

Results

  • Concrete findings:
    • During speech production, participants spent only 7-11% of their time in sustained reading, with most gaze activity focused on monitoring or avoiding distractions.
    • SUMMARY reduced inline regressions and improved readability but increased fixation metrics (e.g., fixation count and duration), indicating unfamiliarity.
    • RAKE was more effective than GP-TSM in guiding reading attention through keyword highlighting.
  • Advantage over baselines:
    • SUMMARY was the most preferred interface for reviewing, significantly outperforming PLAIN, AOC, and GP-TSM in user satisfaction and reading ease.
    • RAKE provided better recall support than GP-TSM, despite both being extractive methods.
  • Experiments / evaluation:
    • Conducted a within-subject experiment with five interfaces, counterbalanced task orders, and a mix of self-spoken and external texts.
    • Metrics included gaze engagement (on/off text, sustained reading, hopping), regressions, and fixation measures (fixation count, duration).
  • Limitations and future work:
    • Limited sample diversity (primarily university students, non-native English speakers).
    • Variability in LLM output stability and API latency.
    • Future work should explore graphical gist representations, personalized highlighting, and broader participant demographics.

Summary

This study investigated visual engagement and reading effort in dictation interfaces using eye-tracking data. It introduced the concept of "Desirable Unfamiliarity," showing that LLM-generated summaries improve readability despite unfamiliar phrasing. RAKE outperformed GP-TSM in guiding attention, while SUMMARY was the most preferred for reviewing. The findings suggest that dictation interfaces should prioritize gist-level representations over verbatim accuracy and reduce distractions during speech production. These insights inform the design of more effective and user-friendly STT systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222346/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791209
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Language Model-Assisted Text Input, AI-Assisted Writing & Text Generation, Eye Tracking & Gaze Interaction
work
Professions
HCI Researchers, University Professors & Researchers, Software Engineers & Developers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers