Interactive Repair of Tables Extracted from PDF Documents on Mobile Devices

Interactive Data VisualizationSoftware Engineers & DevelopersUI/UX DesignersData Scientists & Analysts

PDF documents often contain rich data tables that offer opportunities for dynamic reuse in new interactive applications. We describe a pipeline for extracting, analyzing, and parsing PDF tables based on existing machine learning and rule-based techniques. Implementing and deploying this pipeline on a corpus of 447 documents with 1,171 tables results in only 11 tables that are correctly extracted and parsed. To improve the results of automatic table analysis, we first present a taxonomy of errors that arise in the analysis pipeline and discuss the implications of cascading errors on the user experience. We then contribute a system with two sets of lightweight interaction techniques (gesture and toolbar), for viewing and repairing extraction errors in PDF tables on mobile devices. In an evaluation with 17 users involving both a phone and a tablet, participants effectively repaired common errors in 10 tables, with an average time of about 2 minutes per table.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/5728/2019

AdRecommended

Learn AI Coding at CodeNow

At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2019
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Interactive Data Visualization
work
Professions
Software Engineers & Developers, UI/UX Designers, Data Scientists & Analysts
article
Content Status
Abstract only
hub
Related Papers
10 related papers