Enhancing the Composition Task in Text Entry Studies: Eliciting Difficult Text and Improving Error Rate Calculation
Authors
Document Title
Enhancing the Composition Task in Text Entry Studies: Eliciting Difficult Text and Improving Error Rate Calculation
Document Information
- Domain: Human-Computer Interaction, particularly text entry technologies
- Keywords: text entry, error rate calculation, composition task, natural language input, interface evaluation, autocorrection, smartwatch keyboard, user behavior
Research Background and Problems
- Identified Issues or Challenges:
- In text entry studies, participants typically perform simulated tasks by copying phrases or composing new messages to evaluate interfaces. However, real-world user behavior tends to favor free composition rather than copying phrases.
- Text entry evaluations using static phrase sets fail to adapt to language evolution, and the phrases are often too short (averaging fewer than six words), making it difficult to test the interface's error correction capabilities.
- Composition tasks face challenges: participants may prefer creating simple text, which is insufficient for testing interface performance. Additionally, the lack of reference text affects accurate error rate calculation.
- Significance:
- Composition tasks better reflect real user behavior, aiding in a deeper understanding of text entry interface performance.
- As language evolves and the need for multilingual support grows, novel evaluation methods are essential for text entry interface design.
- Research Motivation and Related Work:
- Related studies indicate that copying tasks are constrained by factors such as user memory load and interaction device limitations. In contrast, composition tasks are more flexible and can accommodate language evolution and diversity.
- Existing methods, such as crowdsourcing proofreading programs, demonstrate certain inaccuracies and require further improvement.
Solution
-
Proposed Methods or Solutions:
- Guiding Participants to Create More Complex Text: Introduce simple guidance strategies to encourage participants to create text that is either easy or difficult to recognize, thus testing the interface's error correction capabilities.
- Accurately Capturing Participants' Intended Text: Obtain reference text through participant input, spontaneous dictation, or crowdsourcing methods, and compare the accuracy of these approaches.
-
Innovations:
- Simple methods to guide participants in adjusting text difficulty
- Comparison of multiple reference text acquisition methods, highlighting their advantages and limitations
- Provision of complex phrases constructed during experiments for future research
-
Implementation Steps and Key Techniques:
- Experiment 1:
- Participants use a smartwatch keyboard to input simple and complex text, followed by entering their intended reference text after the experiment.
- Record character error rate (CER), task time, and correction features (e.g., long-press to lock characters).
- Use crowdsourcing methods to verify the accuracy of reference text.
- Experiment 2:
- Randomly mix complex and simple tasks; participants input text and immediately provide verbal dictation to the experimenter instead of typing it on a computer.
- Recalculate error rates and compare the accuracy of dictation and crowdsourcing methods.
- Experiment 1:
Research Outcomes
-
Specific Results:
- Participants successfully adjusted text difficulty based on instructions.
- Difficult text included more out-of-vocabulary (OOV) words and higher language model complexity (e.g., greater perplexity).
- Across both experiments, complex tasks exhibited higher error rates (CER) and more correction behaviors (e.g., increased use of locking features).
-
Comparison with Existing Solutions:
- Compared to crowdsourcing methods, reference text provided by participants was more accurate.
- Experiments demonstrated that simple instructions effectively increased text complexity and tested interface error correction capabilities, rather than relying on predefined phrase sets.
- Mixing simple and complex tasks, while increasing text complexity, may reduce the "creative flow" of difficult tasks.
-
Experimental or Evaluation Results:
- Results showed that composition tasks have higher external validity (reflecting real behavior) than copying tasks.
- Dictation methods are more suitable than typed references for certain scenarios (e.g., virtual reality experiments) but require real-time supervision by experimenters.
-
Limitations and Future Directions:
- Limitations:
- Current experiments focused on English-proficient university participants, and results may not apply to users with lower language skills or limited keyboard experience.
- Creating complex tasks takes significantly longer (average task time increased by 48%), and methods to enhance user experience are not yet fully developed.
- Future Directions:
- Explore lower-cost or supervision-free methods for reference text acquisition.
- Investigate the applicability of this method to other language input interfaces or specific scenarios (e.g., virtual reality input).
- Sharing complex phrase sets generated from experiments could help improve and evaluate future standard phrase sets.
- Limitations:
Summary
This paper proposes a framework to enhance text composition tasks by guiding text difficulty adjustment and optimizing reference text generation methods, effectively improving the accuracy and representativeness of text entry studies. These findings provide new tools for conducting text entry research closer to real-world usage scenarios and reveal trade-offs among various reference text generation methods.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can users be guided to create texts of varying difficulty to test error-correction capability of text entry interfaces?Category: Input Performance, Accidental Touch Control, and Interaction EfficiencySimilar questionsarrow_forward
- What methods can more accurately obtain users' reference texts?Category: Input Performance, Accidental Touch Control, and Interaction EfficiencySimilar questionsarrow_forward
- How do mixed simple and complex tasks affect participants' text creation experience and validity of input interface evaluation?Category: Input Performance, Accidental Touch Control, and Interaction EfficiencySimilar questionsarrow_forward
Practical Problems
1- Existing text entry research methods fail to reflect real user behavior, especially for complex text input.Category: Input Performance, Accidental Touch Control, and Interaction EfficiencySimilar questionsarrow_forward
- 100%
Timing Matters: How Using LLMs at Different Timings Influences Writers' Perceptions and Ideation Outcomes in AI-Assisted Ideation
CHI '25· Human-LLM Collaboration +1
- 100%
Wordcraft: Story Writing With Large Language Models
IUI '22· Human-LLM Collaboration +1
- 100%
VISAR: A Human-AI Argumentative Writing Assistant with Visual Programming and Rapid Draft Prototyping
UIST '23· Human-LLM Collaboration +1
- 67%
Metaphoria: An Algorithmic Companion for Metaphor Creation
CHI '19· Human-LLM Collaboration +1
- 67%
DiaryMate: Understanding User Perceptions and Experience in Human-AI Collaboration for Personal Journaling
CHI '24· Human-LLM Collaboration +2
- 67%
Understanding Screenwriters' Practices, Attitudes, and Future Expectations in Human-AI Co-Creation
CHI '25· Human-LLM Collaboration +1
- 67%
Letters from Future Self: Augmenting the Letter-Exchange Exercise with LLM-based Agents to Enhance Young Adults' Career Exploration
CHI '25· Human-LLM Collaboration +1
- 67%
Creativity Support in the Age of Large Language Models: An Empirical Study Involving Professional Writers
C&C '24· Human-LLM Collaboration +1
- 67%
Perceptions of Interaction Dynamics in Co-Creative AI: A Comparative Study of Interaction Modalities in Drawcto
C&C '24· Human-LLM Collaboration +1
- 67%
Thoughtful, Confused, or Untrustworthy: How Text Presentation Influences Perceptions of AI Writing Tools
C&C '25· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)