Judging Phishing Under Uncertainty: How Do Users Handle Inaccurate Automated Advice?

Honorable Mention
Privacy by Design & User ControlPrivacy Perception & Decision-MakingCybersecurity EngineersPrivacy Policy Makers

Research Background and Problem Statement

  • What problems or challenges did the authors identify?
    The authors identified that existing email guidance systems have limited effectiveness in helping users recognize phishing emails. This includes the potential of using AI-generated guidance and the impact of inaccurate information within such guidance on user behavior. Some automated systems may incorrectly analyze email content, impairing users' judgment and undermining the system's credibility.

  • Why is this problem important?
    Phishing emails have become a major threat to organizations, often leading to data breaches, account takeovers, and other issues. While current filtering algorithms are effective, some sophisticated phishing emails can bypass defenses and reach users' inboxes, leaving the final judgment to the user. Errors in judgment can result in severe consequences.

  • Research Motivation and Related Work
    The research is motivated by the potential of automated guidance systems to assist in phishing detection and the interest in understanding how humans handle system errors. The authors reviewed existing phishing prevention technologies and their limitations, including attempts to improve user skills through training, the use of warning labels, and the exploration of automated support tools. However, these methods lack highly targeted guidance in "in-the-moment" scenarios. This study aims to address this gap.

Solution

  • What methods or solutions did the authors propose?
    The authors designed and tested three types of guidance conditions: general guidance (Control), perfect guidance tailored to specific emails (Perfect), and realistic guidance containing inaccuracies (Realistic). These conditions provided specific assistance to users during the experiment.

  • What is innovative about this solution?

    1. The study explored the potential of customized, real-time guidance systems to improve users' judgment accuracy and confidence.
    2. It investigated the specific impact of inaccuracies on user behavior, including how different types of guidance content might help or mislead users.
    3. It introduced automatically generated email-based reports that provide explanatory information derived from specific email features (e.g., sender domain, link destination, email subject) to support users' decision-making processes.
  • What are the implementation steps and key technologies used?

    1. The experiment was conducted via an online survey, recruiting 489 participants who were randomly assigned to one of the three guidance conditions.
    2. Participants were shown two rounds of emails: the first round without guidance (R1) and the second round with one of the three types of guidance (R2).
    3. Participants were asked to determine whether each email was a phishing email and to report their confidence levels in their judgments. Feedback on the usability of the reports was also collected after participation.
    4. Statistical analyses were used to compare participants' accuracy and confidence changes between the two rounds and to evaluate the impact of the different guidance categories on these outcomes.

Research Findings

  • What specific results were obtained?

    1. The Perfect guidance condition significantly improved users' judgment accuracy (from 78% to 93%) and confidence (from 3.60 to 4.11).
    2. The Realistic guidance condition, while not improving accuracy, increased users' confidence (from 3.52 to 3.74), indicating that even inaccurate guidance can have a positive impact on users.
    3. The improvement under the Control condition was minimal but still significantly better than having no guidance.
    4. The experiment demonstrated that high-level incorrect advice (e.g., falsely identifying phishing emails as safe) could have a significant negative impact on users' decision-making.
  • What advantages does it have compared to existing solutions?
    Compared to general guidance and URL-based warning methods:

    1. Email-specific customized guidance significantly improved user accuracy, particularly in distinguishing phishing from non-phishing emails.
    2. Automatically generated reports provided explanatory information (e.g., probability scores and evidence analysis), fostering users' trust in the system.
  • What were the experimental or evaluation results?

    1. The experiment showed that email-specific guidance (Perfect condition) was the most effective, significantly outperforming general guidance in terms of accuracy and confidence.
    2. The impact of inaccurate guidance (Realistic condition) depended on the severity of the misleading information. For example, if users recognized incorrect evidence, their confidence decreased; low-confidence classifications also weakened users' trust in the system.
    3. The experiment revealed that users did not blindly trust the system. Even when provided with guidance, users still weighed the suggestions against their own judgment.
  • Limitations and Future Directions

    • Limitations:

      1. The experiment lacked ecological validity, as participants did not replicate real email client interactions in the survey.
      2. The participant sample from the Prolific platform may not fully represent the general user population and was limited to UK-contextualized emails.
      3. The guidance design did not undergo formal pretesting, which may have reduced its utility for users.
      4. The impact of response time and multitasking environments on decision quality was not adequately addressed.
    • Future Directions:

      1. Improve the design of organizational email formats to avoid triggering user misjudgments.
      2. Develop recommendation models based on comprehensive cues to reduce reliance on single indicators (e.g., URLs).
      3. Explore how repeated use of guidance systems can enhance user skills and educational outcomes.
      4. Conduct real-world studies to better understand the long-term impact of automatically generated reports on user behavior.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189664/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714267
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Honorable Mention
group
Authors
5 authors
sell
Subtopics
Privacy by Design & User Control, Privacy Perception & Decision-Making
work
Professions
Cybersecurity Engineers, Privacy Policy Makers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers