Truth or Dare: Understanding and Predicting How Users Lie and Provide Untruthful Data Online
Authors
Title of the Paper
Truth or Dare: Understanding and Predicting How Users Lie and Provide Untruthful Data Online
Paper Information
- Domain: Human-Computer Interaction (HCI), Privacy Protection, and Behavior Prediction
- Keywords: Untruthful Data, Privacy Lies, Privacy Protection Behavior, Prediction Models, False Information, Information Disclosure
Research Background and Problem
-
Problem/Challenge:
- Users frequently provide untruthful data online to protect their privacy, posing a significant challenge for data collectors. Previous studies have primarily focused on "why users lie" and "the impact of lying," but have not delved deeply into "how users lie" or whether such behavior can be predicted.
- There is a lack of systematic understanding of strategies for providing false data and the ability to predict the authenticity of information in specific contexts and for personal data.
-
Significance:
- Individuals providing false data not only affect the effectiveness of their privacy protection but also create issues of data inconsistency or invalidity for data processors. Understanding and predicting users' lying behavior can help design more privacy-friendly data collection mechanisms.
-
Motivation and Related Work:
- Existing research indicates that users tend to lie in privacy-sensitive situations, but the specific strategies and determinants of lying behavior remain underexplored.
- This paper aims to fill this gap by predicting whether users will lie using machine learning models and uncovering common strategies for providing false data.
Solution
-
Proposed Approach:
- Conduct a large-scale empirical study involving over 800 participants, collecting data through Q&A and feedback, and designing a prediction model.
- Use machine learning techniques to predict whether users will provide truthful information in specific scenarios, developing a classifier with high accuracy (89.7%).
- Summarize and categorize four main strategies users employ when lying.
-
Innovations:
- For the first time, the study refines the understanding of user strategies for providing false information, proposing three specific lying methods (invalid information, formatted but entirely false responses, partially truthful answers) and one refusal-to-answer strategy.
- Combines multi-faceted analysis based on contextual variables (comfort, relevance, effort) and personality traits with machine learning.
-
Implementation Steps:
- Study Design: Request users to provide specific personal information in a movie ticket purchase scenario.
- Data Collection: Gather participant feedback on the data items (truthfulness, comfort, etc.) through questionnaires.
- Data Analysis: Employ quantitative methods (machine learning prediction models and regression analysis) and qualitative methods (thematic analysis).
- Model Training: Train a Light Gradient Boosting classifier using partial user feedback to predict data authenticity.
Research Findings
-
Specific Findings:
- Developed a machine learning classifier that predicts whether users will provide truthful information with 89.7% accuracy.
- Regression analysis revealed key factors influencing truthfulness behavior: comfort, relevance, and effort significantly impact behavior.
- Identified three strategies for providing false data and one refusal-to-answer approach.
-
Advantages Over Existing Solutions:
- Improved capability to predict user behavior, providing a basis for improving information interaction design.
- By focusing solely on emotional variables (comfort, relevance, effort), the model avoids using sensitive personal data while effectively predicting behavior.
-
Experimental or Evaluation Results:
- In the training dataset, 85.6% of responses were truthful, and the machine learning model achieved an F1 score of 0.941 on unseen validation data.
- Thematic analysis revealed that completely invalid information was the most common false data strategy, followed by "no response."
-
Limitations and Future Directions:
- Limitations: Some variables (e.g., personality traits) may not be easily collected in real-world applications; the survey design might have subconsciously encouraged lying behavior.
- Future Directions:
- Investigate how data item verification mechanisms influence the generalizability of lying strategy choices.
- Explore the optimization of real data collection and ethical implications through contextual variable control.
- Combine with existing privacy-enhancing technologies (e.g., differential privacy and k-anonymity) to design user-friendly privacy protection tools.
Postscript
This study on user behavior regarding false data provides a crucial foundation for improving applications in the online privacy domain. In particular, the use of machine learning for behavior prediction will significantly impact the field of privacy protection.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What specific deception strategies do users adopt in privacy-sensitive scenarios?Category: Social Platform Safety, Content Governance, and Online HarmSimilar questionsarrow_forward
- Can machine learning models predict users' information truthfulness in specific scenarios?Category: Social Platform Safety, Content Governance, and Online HarmSimilar questionsarrow_forward
- Which contextual variables (comfort, relevance, effort, etc.) most affect users' deceptive behavior?Category: Social Platform Safety, Content Governance, and Online HarmSimilar questionsarrow_forward
Practical Problems
1- To protect privacy, users often provide false personal information online.Category: Social Platform Safety, Content Governance, and Online HarmSimilar questionsarrow_forward
- 80%
A Psychometric Scale to Measure Individuals' Value of Other People's Privacy (VOPP)
CHI '23· AI Ethics, Fairness & Accountability +2
- 67%
Exploring What People Need to Know to be AI Literate: Tailoring for a Diversity of AI Roles and Responsibilities
CHI '25· Explainable AI (XAI) +2
- 67%
When Feasibility of Fairness Audits Relies on Willingness to Share Data: Examining User Acceptance of Multi-Party Computation Protocols for Fairness Monitoring
CHI '26· AI Ethics, Fairness & Accountability +2
- 67%
Do Citizens Agree with the EU AI Act? Public Perspectives on Risk and Regulation of AI Systems
CHI '26· AI Ethics, Fairness & Accountability +2
- 60%
I don't need an expert! Making URL phishing features human comprehensible
CHI '21· Algorithmic Transparency & Auditability +1
- 60%
Assessing MyData Scenarios: Ethics, Concerns, and the Promise
CHI '21· AI Ethics, Fairness & Accountability +1
- 60%
Toggles, Dollar Signs, and Triangles: How to (In)Effectively Convey Privacy Choices
CHI '21· Privacy by Design & User Control +1
- 60%
Covert Embodied Choice: Decision-Making and the Limits of Privacy Under Biometric Surveillance
CHI '21· Privacy by Design & User Control +1
- 60%
"Okay, whatever": An Evaluation of Cookie Consent Interfaces
CHI '22· Privacy Perception & Decision-Making +1
- 60%
“Our Users' Privacy is Paramount to Us”: A Discourse Analysis of How Period and Fertility Tracking App Companies Address the Roe v Wade Overturn
CHI '24· Privacy by Design & User Control +1
Based on Jaccard similarity of research subtopics & professions (≥60%)