Using Open Data to Automatically Generate Localized Analogies
Interactive Data VisualizationData StorytellingComputational Methods in HCIHCI ResearchersStatisticians & Data Scientists
Document Title
Automatically Generating Localized Analogies Using Open Data
Document Information
- Subject Area: Human-Computer Interaction (HCI), Improving Digital Analogies Using Open Data
- Keywords: analogy, numerical understanding, measurement, crowdsourcing, open data, localization, Wikipedia, user adaptation
Research Background and Problem Statement
-
Identified Problems:
- Numerical information is ubiquitous in news and statistics, yet many people struggle to understand or identify incorrect figures, even in critical areas such as natural disasters, policy, or health.
- Using analogies (i.e., "perspectives") to transform complex numbers into easily comprehensible forms has proven to enhance readers' understanding and memory, but current methods face challenges in scalability, localization, and multicultural adaptation.
- Existing analogy generation methods rely on crowdsourcing or domain-specific models, which struggle to reach broad audiences and effectively adapt to different countries and cultures.
-
Significance of Research:
- Accurate understanding of numerical information is crucial for voters, consumers, and policymakers.
- In the context of globalization, conveying complex statistical information to audiences with diverse cultures and languages is increasingly important.
-
Motivation and Related Work:
- Current methods often rely on time-consuming crowdsourcing, which is limited by the characteristics of user groups from specific countries or languages, resulting in restricted localization effects.
- This study aims to develop a scalable and international solution by leveraging automation and open data resources (e.g., Wikidata and Wikipedia).
Solution
-
Proposed Method:
- Construct a reference database containing 39,532 fundamental objects by extracting measurement information such as length, height, area, and mass from Wikidata.
- Design a model that uses various open data signals (e.g., search volume, text embeddings, N-gram statistics, and Wiki page traffic) to build proxy features for predicting the suitability of analogies.
- Simplify the model design, demonstrating that a model based solely on Wikipedia traffic can deliver efficient performance, thereby reducing computational and modeling costs.
- Evaluate the model: test its performance in adapting to new measurement types (e.g., demographics) and different cultural contexts (e.g., the U.S. and France).
-
Innovations:
- Utilizing Wikipedia traffic as a proxy feature circumvents the need for manual intervention and complex models, significantly reducing the difficulty and cost of generating adaptive analogies.
- The model does not require retraining for new domains or cultures, making it easier to scale.
-
Implementation Steps:
- Data Collection: Use the SPARQL API of Wikidata to retrieve measurement information and links to multiple languages.
- Data Processing: Clean and integrate redundant, ambiguous, and impractical objects to establish a high-quality candidate set.
- Model Design: Use signals linking reference objects to users as input features to validate their predictive power for analogy suitability.
- User Validation: Collect subjective evaluations of analogies through surveys and use them as benchmarks for assessing model performance.
Research Results
Specific Outcomes
- The model can generate high-quality analogies solely based on Wikipedia page traffic signals, which are generally useful and easy for users to understand.
- The model performs well in untrained domains such as demographics and new measurement types (Spearman correlation coefficient ρ up to 0.83).
- The model also demonstrates strong cultural adaptability, generating different and more acceptable analogies for U.S. and French audiences.
Advantages Over Existing Solutions
- Avoids the high cost and cultural singularity of crowdsourcing by achieving automation and scalability through open data resources.
- Adapts to new measurement types and audiences from different languages/cultures without requiring model retraining.
- Leverages Wikipedia, a cross-cultural and global open-source platform, to enhance global applicability.
Experimental or Evaluation Results
- Model Performance:
- Analogy generation for existing measurement types: ρ = 0.77 (overall).
- Outstanding adaptability in new measurement types (e.g., demographics): ρ = 0.83.
- Cultural Adaptability:
- U.S. users preferred culturally matched analogies 80% of the time.
- In experiments, French users favored localized perspectives 59% of the time.
- User Estimation Experiments:
- U.S. users using localized analogies had significantly lower estimation errors for measurement values compared to those using non-localized analogies (error magnitude differences of 4x vs. 10x).
Limitations and Future Directions
-
Limitations:
- Open Data does not include all potentially useful common objects (e.g., beds, microwaves).
- Wikipedia traffic signals exhibit uneven language coverage and insufficient granularity in cross-cultural contexts.
- Low-quality entries in knowledge bases may interfere with the model, such as ambiguous entities (e.g., unnamed roads or specific sculptures) that may be irrelevant.
-
Future Work:
- Integrate crowdsourcing to supplement missing information in Wikidata or Wikipedia.
- Provide multiple analogy options instead of a single reference to enhance user adaptability.
- Explore fine-tuning within regions of the same country, such as differentiating between U.S. and U.K. cultural adaptations.
- Validate the model across more cultural and linguistic contexts to ensure universal applicability and fairness.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can open data enable automatic analogy generation for complex numerical information and improve users' numerical comprehension?Category: News Exposure, Information Literacy, and Fact-Checking AnalysisSimilar questionsarrow_forward
- How can automatic analogy generation ensure adaptability in multicultural and multilingual contexts?Category: News Exposure, Information Literacy, and Fact-Checking AnalysisSimilar questionsarrow_forward
- Can models based solely on Wikipedia traffic data efficiently generate culturally appropriate analogies?Category: News Exposure, Information Literacy, and Fact-Checking AnalysisSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Ordinary users struggle to understand complex numerical information in news or statistics.Category: News Exposure, Information Literacy, and Fact-Checking AnalysisSimilar questionsarrow_forward
- 80%
ConceptScope: Organizing and Visualizing Knowledge in Documents based on Domain Ontology
CHI '21· Interactive Data Visualization +1
- 67%
CrossData: Leveraging Text-Data Connections for Authoring Data Documents
CHI '22· Interactive Data Visualization +1
- 67%
Idyll Studio: A Structured Editor for Authoring Interactive & Data-Driven Articles
UIST '21· Interactive Data Visualization +1
- 60%
From Detectables to Inspectables: Understanding Qualitative Analysis of Audiovisual Data
CHI '21· Interactive Data Visualization +1
- 60%
Interaction Illustration Taxonomy: Classification of Styles and Techniques for Visually Representing Interaction Scenarios
CHI '21· Interactive Data Visualization +1
- 60%
Interactive Context-Preserving Color Highlighting for Multiclass Scatterplots
CHI '23· Interactive Data Visualization
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642638
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Interactive Data Visualization, Data Storytelling, Computational Methods in HCI
work
Professions
HCI Researchers, Statisticians & Data Scientists
article
Content Status
Full text indexed
hub
Related Papers
6 related papers