Desirable Unfamiliarity: Insights from Eye Movements on Engagement and Readability of Dictation Interfaces
Authors
Paper Title
Desirable Unfamiliarity: Insights from Eye Movements on Engagement and Readability of Dictation Interfaces
Publication Info
- Topic area: Eye-tracking analysis of dictation interfaces for speech-to-text systems.
- Keywords: Dictation interfaces, speech-to-text, eye-tracking, reading behavior, LLM-generated summaries, keyword highlighting, user engagement, cognitive load, text readability, transcription systems.
Background and Problem
- Problem / challenge: Raw speech-to-text (STT) transcripts are often hard to read due to disfluencies, recognition errors, and verbosity. While LLM-based corrections can improve readability, they risk altering the original meaning or causing distractions during speech production.
- Significance: Improving the readability and usability of dictation interfaces is critical for their adoption in personal, professional, and educational contexts.
- Motivation and related work: Existing research has focused on transcription accuracy and disfluency correction but has not evaluated the effectiveness of readability-enhancing solutions like keyword highlighting or LLM-generated summaries. Eye-tracking studies on reading behavior in dictation interfaces remain unexplored.
Solution
- Proposed approach: Conduct an eye-tracking study to evaluate five dictation interfaces: PLAIN, AOC, RAKE, GP-TSM, and SUMMARY, each differing in fidelity to original speech, review aids, and interface responsiveness.
- Novelty:
- Introduced the concept of "Desirable Unfamiliarity," where LLM-generated summaries improve readability despite introducing unfamiliar words.
- Compared extractive (RAKE, GP-TSM) and abstractive (SUMMARY) methods for aiding text review in dictation interfaces.
- Analyzed eye-tracking data to understand visual engagement and reading effort during speech production and review.
- Provided actionable design implications for future dictation interfaces.
- Procedure and key techniques:
- Participants (N=20) composed and reviewed social media posts using five interfaces while eye movements were tracked.
- Interfaces included real-time transcription (PLAIN), periodic corrections (AOC), keyword highlighting (RAKE), grammar-preserving highlights (GP-TSM), and LLM-generated summaries (SUMMARY).
- Measured gaze engagement, regressions, and fixation metrics during production and review phases.
- Conducted qualitative analysis through surveys and interviews.
Results
- Concrete findings:
- During speech production, participants spent only 7-11% of their time in sustained reading, with most gaze activity focused on monitoring or avoiding distractions.
- SUMMARY reduced inline regressions and improved readability but increased fixation metrics (e.g., fixation count and duration), indicating unfamiliarity.
- RAKE was more effective than GP-TSM in guiding reading attention through keyword highlighting.
- Advantage over baselines:
- SUMMARY was the most preferred interface for reviewing, significantly outperforming PLAIN, AOC, and GP-TSM in user satisfaction and reading ease.
- RAKE provided better recall support than GP-TSM, despite both being extractive methods.
- Experiments / evaluation:
- Conducted a within-subject experiment with five interfaces, counterbalanced task orders, and a mix of self-spoken and external texts.
- Metrics included gaze engagement (on/off text, sustained reading, hopping), regressions, and fixation measures (fixation count, duration).
- Limitations and future work:
- Limited sample diversity (primarily university students, non-native English speakers).
- Variability in LLM output stability and API latency.
- Future work should explore graphical gist representations, personalized highlighting, and broader participant demographics.
Summary
This study investigated visual engagement and reading effort in dictation interfaces using eye-tracking data. It introduced the concept of "Desirable Unfamiliarity," showing that LLM-generated summaries improve readability despite unfamiliar phrasing. RAKE outperformed GP-TSM in guiding attention, while SUMMARY was the most preferred for reviewing. The findings suggest that dictation interfaces should prioritize gist-level representations over verbatim accuracy and reduce distractions during speech production. These insights inform the design of more effective and user-friendly STT systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)