Lending a Hand: The Effectiveness of Support Systems in Assisting Users to Detect Phishing Attacks
Authors
Paper Title
Lending a Hand: The Effectiveness of Support Systems in Assisting Users to Detect Phishing Attacks
Publication Info
- Topic area: Evaluation of anti-phishing support systems in realistic email classification scenarios.
- Keywords: Phishing detection, support systems, user behavior, email security, tooltip, warning banner, external marker, spam label, user experience.
Background and Problem
- Problem / challenge: Despite the prevalence of anti-phishing support systems in email clients, their effectiveness in helping users detect phishing emails remains unclear. Prior studies have not comprehensively evaluated multiple support systems in realistic scenarios, especially considering false positives and false negatives.
- Significance: Phishing attacks are a major cybersecurity threat, responsible for 41% of account compromises in Q1 2025. With advancements in generative AI, phishing emails are becoming harder to detect. Effective support systems could significantly reduce phishing-related risks.
- Motivation and related work: Previous research has focused on training users or warning systems but has not comparatively evaluated multiple support systems in realistic, controlled environments. This study addresses this gap by testing four support systems and their impact on phishing detection.
Solution
- Proposed approach: Development of an interactive tool simulating a realistic email classification scenario, where participants evaluate emails with or without support systems.
- Novelty:
- Comparative evaluation of four real-world and academic support systems: external marker, warning banner, spam label, and tooltip.
- Inclusion of false positives and false negatives in the evaluation to mimic real-world imperfections.
- Analysis of contextual factors (e.g., contact lists, calendars) and user behavior in phishing detection.
- Collection of subjective user feedback and user experience metrics (UEQ-S).
- Procedure and key techniques:
- Participants (n=453) role-played as HR employees, classifying 18 emails (3 phishing, 15 legitimate) using an interactive tool.
- Support systems were randomly assigned to groups: external marker (n=99), warning banner (n=87), spam label (n=85), tooltip (n=91), and a control group (n=91).
- Behavioral data (e.g., link hovering, time spent) and survey responses were collected.
- Statistical analysis included Kruskal-Wallis tests, Mann-Whitney U tests, and logistic regression.
Results
- Concrete findings:
- Participants correctly classified 83.9%–88.3% of emails on average, with no significant improvement from support systems compared to the control group.
- Contextual factors, such as using the contact list (+93.2% correct classification) and hovering over links (+91.2%), had a stronger impact than support systems.
- False positives and false negatives in support systems had minimal influence on subsequent classifications.
- Advantage over baselines:
- The warning banner slightly outperformed other systems in specific scenarios but did not significantly improve overall classification rates.
- Tooltip usage led to a slight decline in performance for certain legitimate emails.
- Experiments / evaluation:
- Tested emails included varying difficulty levels (easy, medium, difficult).
- Metrics: classification accuracy, user perception (survey), and user experience (UEQ-S).
- Participants appreciated support systems, rating them positively on pragmatic quality (mean scores: banner 1.7, external marker 1.6, tooltip 1.0, label 1.3).
- Limitations and future work:
- Study conducted in a controlled, simulated environment; real-world applicability may differ.
- Limited phishing email types and an unbalanced ratio of phishing to legitimate emails.
- Future work should test support systems in real workplaces with diverse participant samples and realistic email ratios.
Summary
This study evaluated the effectiveness of four anti-phishing support systems in a controlled email classification scenario with 453 participants. Results showed no significant improvement in phishing detection from support systems compared to a control group, with contextual factors like contact lists and link hovering playing a more critical role. Despite limited practical impact, participants perceived the systems as helpful and rated them positively in user experience surveys. Future research should explore real-world implementations and refine support systems to address false positives, unclear assessments, and user trust.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)