AnnotateGPT: Designing Human–AI Collaboration in Pen-Based Document Annotation
Authors
Paper Title
AnnotateGPT: Designing Human–AI Collaboration in Pen-Based Document Annotation
Publication Info
- Topic area: Human–AI collaboration in educational feedback through pen-based annotation systems.
- Keywords: Human–AI collaboration, pen-based annotation, large language models, educational feedback, document annotation, purpose inference, feedback generation, teacher support, annotation clustering, interactive AI systems.
Background and Problem
- Problem / challenge: Providing high-quality feedback on writing is cognitively demanding and time-consuming for educators. Traditional handwritten annotations are valued for their personal tone but suffer from issues like poor legibility, time constraints, and inconsistent quality. Existing digital tools often reduce annotations to static marks, missing opportunities for interactive, AI-augmented feedback.
- Significance: Addressing these challenges could enhance the efficiency and quality of feedback, particularly in educational contexts where timely, constructive feedback is critical for student learning.
- Motivation and related work: Prior systems like XLibris and Metatation explored annotation-driven augmentation but were domain-specific and lacked purpose inference. Mainstream tools like Grammarly provide sentence-level corrections but fail to integrate with personalized annotation workflows. AnnotateGPT builds on these gaps by leveraging LLMs to interpret handwritten annotations and generate contextually relevant feedback.
Solution
- Proposed approach: AnnotateGPT, a pen-based annotation system that uses large language models (LLMs) to infer the purpose of handwritten annotations and generate expanded feedback across documents.
- Novelty:
- Integration of LLMs to classify annotation purposes and propagate feedback throughout a document.
- Use of pen annotations as implicit interaction signals for AI collaboration.
- Design of workflows that balance human agency and AI scalability in feedback generation.
- Procedure and key techniques:
- Stroke clustering: Hierarchical agglomerative clustering groups pen strokes into annotation clusters based on spatiotemporal distance.
- Stroke classification and text extraction: Pen strokes are classified into types (e.g., highlighting, circling) to extract text for LLM input.
- Purpose inference: LLMs propose four possible purposes for each annotation, guided by structured prompts and prior annotation history.
- Feedback generation: LLMs generate context-specific annotations, targeting sentences and words based on user-selected purposes.
- Interactive interface: Features include assistant markers, tooltips for feedback verification, and a specialized scrollbar for navigation.
Results
- Concrete findings:
- AnnotateGPT inferred annotation purposes correctly in 63% of cases, with explicit annotations achieving 100% accuracy.
- Participants made significantly fewer annotations with AnnotateGPT (M=4 per question) compared to the baseline (M=7 per question).
- AnnotateGPT generated more annotations overall, with 47 accepted, 19 rejected, and 8 marked as helpful on average per participant.
- Advantage over baselines:
- AnnotateGPT reduced physical demand (p < 0.05) and enabled faster, more consistent feedback generation.
- Participants reported that AnnotateGPT provided broader, more detailed feedback compared to the baseline.
- Experiments / evaluation:
- A user study with 12 pre-service teachers evaluated AnnotateGPT against a baseline digital annotation tool.
- Data included annotations, system logs, questionnaires, and interviews.
- Metrics: annotation density, duration, usability (SUS), workload (NASA-TLX), and feedback quality.
- Limitations and future work:
- Small, homogeneous participant pool limits generalizability.
- Study focused on teacher experiences, not student outcomes.
- Errors in purpose inference and feedback generation occurred, particularly with telegraphic annotations.
- Future work should explore larger, diverse samples, long-term impacts, and multimodal interactions.
Summary
AnnotateGPT leverages LLMs to transform handwritten annotations into interactive, AI-augmented feedback tools. By inferring annotation purposes and generating expanded feedback, it enhances efficiency and quality in educational contexts. A user study demonstrated its ability to reduce workload, provide broader feedback, and maintain teacher agency. However, challenges like misclassification of telegraphic annotations and limited generalizability highlight areas for improvement. AnnotateGPT offers a promising framework for human–AI collaboration in annotation, with potential applications beyond education.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 63%
From Crafting Text to Crafting Thought: Grounding AI Writing Support to Writing Center Pedagogy
CHI '26· Human-LLM Collaboration +2
- 63%
Co-Designing with Algorithms: Unpacking the Complex Role of GenAI in Interactive System Design Education
DIS '25· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)