PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs
The document contains substantial unannotated data, necessitating extensive manual labeling efforts. To address this issue, we introduce PDFChatAnnotator, a human-LLM collaborative tool to collect multi-modal data from PDF catalogs. Initially, PDFChatAnnotator automatically employs our proposed multi-modal binding rules to link related data from different modalities and harnesses the information extraction capabilities of large language models (LLMs) to extract specific information from text descriptions. Furthermore, the tool empowers users to guide and refine the LLM's annotations. During the annotation process, users can influence the LLM through multiple rounds of communication and example establishment via the provided interfaces. To assess the effectiveness of PDFChatAnnotator's techniques, we conducted a technical evaluation using three catalogs with typical layouts as experimental data. The results showed that all accuracy rates for multi-modal binding exceeded 90%, and both the proposed "example establishment" and "interactive adjustment of requirements" contributed to enhanced accuracy rates.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 80%
Generating Automatic Feedback on UI Mockups with Large Language Models
CHI '24· Human-LLM Collaboration +1
- 80%
DynEx: Dynamic Code Synthesis with Structured Design Exploration for Accelerated Exploratory Programming
CHI '25· Human-LLM Collaboration +1
- 80%
Towards Rapid Interactive Machine Learning: Evaluating Tradeoffs of Classification without Representation
IUI '19· Human-LLM Collaboration +1
- 80%
StoryEnsemble: Enabling Dynamic Exploration & Iteration in the Design Process with AI and Forward-Backward Propagation
UIST '25· Human-LLM Collaboration +1
- 71%
Botender: Supporting Communities in Collaboratively Designing AI Agents through Case-Based Provocations
CHI '26· Human-LLM Collaboration +3
- 71%
The Invisible Mentor: Inferring User Actions from Screen Recordings to Recommend Better Workflows
CHI '26· Human-LLM Collaboration +3
- 67%
May AI? Design Ideation with Cooperative Contextual Bandits
CHI '19· Generative AI (Text, Image, Music, Video) +2
- 67%
An Analytic Model for Time Efficient Personal Hierarchies
CHI '19· User Research Methods (Interviews, Surveys, Observation) +1
- 67%
Swap: A Replacement-based Text Revision Technique for Mobile Devices
CHI '20· Context-Aware Computing +1
- 67%
The Image of the Interface: How People Use Landmarks to Develop Spatial Memory of Commands in Graphical Interfaces
CHI '21· Visualization Perception & Cognition +1
Based on Jaccard similarity of research subtopics & professions (≥60%)