DataPilot: Utilizing Quality and Usage Information for Subset Selection during Visual Data Preparation
Authors
Title of the Paper
DataPilot: Utilizing Quality and Usage Information for Subset Selection during Visual Data Preparation
Bibliographic Information
- Research Domain: Data Visualization, Data Preparation, Data Analysis Tools
- Keywords: Data Quality, Data Usage, Subset Selection, Data Preparation, Visualization, Visual Data Analysis, Design Research
Research Background and Problem Statement
-
What problems or challenges did the authors identify?
- Selecting relevant data subsets is challenging when dealing with large-scale, unfamiliar datasets, as users may overlook important attributes or select irrelevant data due to improper handling.
- Data quality issues (e.g., missing data, inaccuracies) are pervasive and can amplify errors in analysis results.
- The potential of utilizing other users' historical analysis trajectories ("usage information") remains underdeveloped.
-
Why is this problem important?
- Data quality directly impacts analysis outcomes; low-quality data leads to wasted resources and erroneous decision-making.
- Current visualization and data analysis tools lack user-guided designs to assist in selecting data subsets effectively.
-
Research Motivation and Related Work
- Through design research, the authors observed a strong demand among data users for high-quality data, yet existing tools provide insufficient visualization support for data quality and usage information.
- While current tools and literature address data quality to some extent, research on modeling and presenting data usage information is limited.
Solution
-
What methods or solutions did the authors propose?
- Introduced a visualization-based data preparation and analysis tool named DataPilot, which integrates data quality and usage information to guide users in selecting data subsets.
- Developed a three-dimensional model for data quality: completeness, correctness, and objectivity; and a three-dimensional model for data usage information: in-subsets, in-filters, and in-visualizations.
- Designed an interactive visualization interface to provide effective guidance based on these dimensions.
-
What are the innovative aspects of this solution?
- Combined "quality" and "usage" information using data logs and metadata, presenting them as standardized scores.
- Incorporated interactive prompts (e.g., color cues and sorting controls) into the data analysis tool to streamline the subset selection process.
- Implemented an innovative "multi-step workflow" design that directly integrates data preparation with visual analysis.
-
What are the implementation steps and key technologies used?
- Data Modeling:
- Defined and calculated quality scores for data attributes, including completeness, correctness, and objectivity.
- Built historical usage information from user logs, such as frequency of usage.
- User Interface Design:
- Developed a three-step workflow UI: raw data inspection, subset selection review, and dashboard creation.
- Provided visual interaction features such as filtering and sorting.
- User Research Testing:
- Conducted experiments comparing tool performance under conditions with only quality information, basic settings, and both quality and usage information.
- Data Modeling:
Research Outcomes
-
What specific outcomes were achieved?
- Compared to traditional methods, users of DataPilot could select smaller yet more relevant data subsets, improving the quality and efficiency of subset selection.
- Users demonstrated higher success rates and confidence when selecting data subsets.
- Survey results revealed that users responded more positively to data quality dimensions than to usage information.
-
What advantages does it have over existing solutions?
- Compared to other data visualization tools (e.g., Tableau), DataPilot introduces intuitive representations of data quality and usage information.
- User studies confirmed that combining feedback from both types of information is more effective in completing data preparation tasks.
- The three-step workflow design reduces the need for users to frequently switch between task interfaces.
-
What were the experimental or evaluation results?
- Under three user conditions (no data information, quality information only, and both quality and usage information), users employing DataPilot selected datasets with higher quality metrics (e.g., fewer missing values, higher relevance).
- Results indicated that users prioritized completeness as their preferred strategy for subset filtering.
-
Limitations and Future Directions
- Limitations:
- The current tool supports only tabular data and is not applicable to unstructured or multidimensional data analysis.
- The modeling rules for data quality dimensions and usage information are relatively fixed, limiting cross-domain applicability.
- Some users expressed caution about relying on historical usage information, citing concerns over guided conformity.
- Future Directions:
- Design more flexible configuration options, allowing users to customize weights or rules for data quality dimensions.
- Extend the tool's applicability to machine learning and text data processing domains.
- Explore additional interactive mechanisms in the data preparation process, such as dynamic recommendations or novel user guidance methods.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can data quality (completeness, correctness, and objectivity) and historical usage information improve subset selection efficiency in large datasets?Category: Context-Aware Sampling and Low-Disruption NotificationsSimilar questionsarrow_forward
- How can displaying data quality and usage information in data analysis tools help users filter higher-quality data subsets?Category: Context-Aware Sampling and Low-Disruption NotificationsSimilar questionsarrow_forward
- Can integrating multi-step workflows into data preparation and visual analysis improve user task efficiency?Category: Context-Aware Sampling and Low-Disruption NotificationsSimilar questionsarrow_forward
Practical Problems
1- Users struggle to select high-quality analysis subsets from large datasets.Category: Context-Aware Sampling and Low-Disruption NotificationsSimilar questionsarrow_forward
- 100%
PriorWeaver: Prior Elicitation via Iterative Dataset Construction
CHI '26· Interactive Data Visualization +1
- 100%
Taking Truncation to Task: A Task-Based Exploration of Axis Truncation in Bar Charts
CHI '26· Interactive Data Visualization +1
- 100%
D-MO: Depth from Motion and Occlusion as a Visual Channel for Information Visualization
CHI '26· Interactive Data Visualization +1
- 80%
Effects of Point Size and Opacity Adjustments in Scatterplots
CHI '24· Interactive Data Visualization +1
- 80%
Lost in Magnitudes: Exploring Visualization Designs for Large Value Ranges
CHI '25· Interactive Data Visualization +2
- 80%
AVEC: An Assessment of Visual Encoding Ability in Visualization Construction
CHI '25· Interactive Data Visualization +1
- 80%
Examining Interpretation Strategies for Multiple Forecast Visualizations with Two and Four Forecasts
CHI '26· Interactive Data Visualization +2
- 80%
Revealing the Gap: Visual Comparison of Large-Scale Datasets via Multi-Scale Density Difference Map
CHI '26· Interactive Data Visualization +1
- 80%
Studying the Separability of Visual Channel Pairs in Symbol Maps
CHI '26· Interactive Data Visualization +2
- 80%
Redundant is Not Redundant: Automating Efficient Categorical Palettes Design Unifying Color & Shape Encodings with CatPAW
CHI '26· Interactive Data Visualization +1
Based on Jaccard similarity of research subtopics & professions (≥60%)