Truth or Dare: Understanding and Predicting How Users Lie and Provide Untruthful Data Online

AI Ethics, Fairness & AccountabilityPrivacy Perception & Decision-MakingPrivacy Policy MakersHCI Researchers

Title of the Paper

Truth or Dare: Understanding and Predicting How Users Lie and Provide Untruthful Data Online

Paper Information

  • Domain: Human-Computer Interaction (HCI), Privacy Protection, and Behavior Prediction
  • Keywords: Untruthful Data, Privacy Lies, Privacy Protection Behavior, Prediction Models, False Information, Information Disclosure

Research Background and Problem

  • Problem/Challenge:

    • Users frequently provide untruthful data online to protect their privacy, posing a significant challenge for data collectors. Previous studies have primarily focused on "why users lie" and "the impact of lying," but have not delved deeply into "how users lie" or whether such behavior can be predicted.
    • There is a lack of systematic understanding of strategies for providing false data and the ability to predict the authenticity of information in specific contexts and for personal data.
  • Significance:

    • Individuals providing false data not only affect the effectiveness of their privacy protection but also create issues of data inconsistency or invalidity for data processors. Understanding and predicting users' lying behavior can help design more privacy-friendly data collection mechanisms.
  • Motivation and Related Work:

    • Existing research indicates that users tend to lie in privacy-sensitive situations, but the specific strategies and determinants of lying behavior remain underexplored.
    • This paper aims to fill this gap by predicting whether users will lie using machine learning models and uncovering common strategies for providing false data.

Solution

  • Proposed Approach:

    • Conduct a large-scale empirical study involving over 800 participants, collecting data through Q&A and feedback, and designing a prediction model.
    • Use machine learning techniques to predict whether users will provide truthful information in specific scenarios, developing a classifier with high accuracy (89.7%).
    • Summarize and categorize four main strategies users employ when lying.
  • Innovations:

    • For the first time, the study refines the understanding of user strategies for providing false information, proposing three specific lying methods (invalid information, formatted but entirely false responses, partially truthful answers) and one refusal-to-answer strategy.
    • Combines multi-faceted analysis based on contextual variables (comfort, relevance, effort) and personality traits with machine learning.
  • Implementation Steps:

    1. Study Design: Request users to provide specific personal information in a movie ticket purchase scenario.
    2. Data Collection: Gather participant feedback on the data items (truthfulness, comfort, etc.) through questionnaires.
    3. Data Analysis: Employ quantitative methods (machine learning prediction models and regression analysis) and qualitative methods (thematic analysis).
    4. Model Training: Train a Light Gradient Boosting classifier using partial user feedback to predict data authenticity.

Research Findings

  • Specific Findings:

    • Developed a machine learning classifier that predicts whether users will provide truthful information with 89.7% accuracy.
    • Regression analysis revealed key factors influencing truthfulness behavior: comfort, relevance, and effort significantly impact behavior.
    • Identified three strategies for providing false data and one refusal-to-answer approach.
  • Advantages Over Existing Solutions:

    • Improved capability to predict user behavior, providing a basis for improving information interaction design.
    • By focusing solely on emotional variables (comfort, relevance, effort), the model avoids using sensitive personal data while effectively predicting behavior.
  • Experimental or Evaluation Results:

    • In the training dataset, 85.6% of responses were truthful, and the machine learning model achieved an F1 score of 0.941 on unseen validation data.
    • Thematic analysis revealed that completely invalid information was the most common false data strategy, followed by "no response."
  • Limitations and Future Directions:

    • Limitations: Some variables (e.g., personality traits) may not be easily collected in real-world applications; the survey design might have subconsciously encouraged lying behavior.
    • Future Directions:
      • Investigate how data item verification mechanisms influence the generalizability of lying strategy choices.
      • Explore the optimization of real data collection and ethical implications through contextual variable control.
      • Combine with existing privacy-enhancing technologies (e.g., differential privacy and k-anonymity) to design user-friendly privacy protection tools.

Postscript

This study on user behavior regarding false data provides a crucial foundation for improving applications in the online privacy domain. In particular, the use of machine learning for behavior prediction will significantly impact the field of privacy protection.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47406/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445625
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
AI Ethics, Fairness & Accountability, Privacy Perception & Decision-Making
work
Professions
Privacy Policy Makers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers