Deception at Scale: Deceptive Designs in 1K LLM-Generated E-Commerce Components

AI Ethics, Fairness & AccountabilityDark Patterns RecognitionHuman-LLM CollaborationAI/ML Researchers & EngineersUI/UX DesignersE-Commerce Platform Operators

Paper Title

Deception at Scale: Deceptive Designs in 1K LLM-Generated E-Commerce Components

Publication Info

  • Topic area: Large Language Models (LLMs) and deceptive design patterns in e-commerce interfaces.
  • Keywords: LLMs, deceptive designs, dark patterns, e-commerce, interface interference, user-centered design, business incentives, human values, prompt engineering, front-end code.

Background and Problem

  • Problem / challenge: LLM-generated front-end code often embeds deceptive designs, potentially scaling harmful practices unintentionally.
  • Significance: Deceptive designs can harm users financially, invade privacy, and create cognitive burdens, and their automated generation risks widespread deployment without oversight.
  • Motivation and related work: Prior studies have documented deceptive designs in human-designed interfaces and found initial evidence of dark patterns in LLM-generated code. However, systematic, large-scale investigations into the frequency, strategies, and mitigation of deceptive designs in LLM outputs were lacking.

Solution

  • Proposed approach: A large-scale audit of LLM-generated e-commerce components to identify deceptive designs, assess influencing factors, and test strategies to reduce their frequency.
  • Novelty:
    1. First large-scale audit of deceptive designs in LLM-generated front-end code.
    2. Introduction of design rationales to infer intent and interaction assumptions.
    3. Testing of stakeholder interest prompts and user-centered strategies to mitigate deceptive designs.
    4. Development of a dataset of 1,296 annotated e-commerce components.
  • Procedure and key techniques:
    • Study 1: Generated 1,080 components across 15 e-commerce types using four LLMs under three stakeholder interest conditions (business, user, baseline). Annotated designs using Gray et al.’s taxonomy.
    • Study 2: Tested three new user-centered prompt strategies (direct mitigation, deception definition, human values) on six components using two LLMs. Compared outputs for reductions in deceptive designs.

Results

  • Concrete findings:
    • 55.8% of components contained at least one deceptive design, and 30.6% contained two or more.
    • Interface interference was the dominant strategy (63.2% of instances), followed by forced action (15.3%) and social engineering (14.0%).
    • Business interest prompts increased deceptive designs by 15.8 percentage points, while user interest prompts reduced them by 5.8 percentage points.
    • Human value prompts reduced deceptive designs by 15.3 percentage points compared to baseline.
  • Advantage over baselines:
    • DeepSeek-V3 produced significantly fewer deceptive designs (46.3%) compared to Gemini 2.5 Pro (60.7%), Grok 3 Beta (60.4%), and GPT-4.1 (55.9%).
    • Human value prompts were the most effective mitigation strategy, outperforming direct mitigation and deception definition.
  • Experiments / evaluation:
    • Study 1: Factorial design testing 15 components, 4 LLMs, and 3 stakeholder interest conditions. Manual annotation of 1,080 outputs.
    • Study 2: Stratified sampling of 6 components and 2 models under 3 new user-centered strategies. Manual annotation of 216 outputs.
  • Limitations and future work:
    • Focused only on e-commerce components; other domains may exhibit different patterns.
    • Static components excluded JavaScript-based dynamic behaviors.
    • Some annotation disagreements highlight the need for refined criteria and further study of controversial cases.

Summary

This paper presents the first large-scale audit of deceptive designs in LLM-generated e-commerce components, revealing that over half of outputs contained at least one dark pattern. The study identified interface interference as the dominant strategy and showed that business interest prompts significantly increased deceptive designs, while human value prompts were the most effective at reducing them. By introducing design rationales and testing mitigation strategies, the paper highlights systemic risks in LLM-assisted code generation and offers actionable recommendations for developers and providers to embed safeguards and human values into future systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222135/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791063
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
AI Ethics, Fairness & Accountability, Dark Patterns Recognition, Human-LLM Collaboration
work
Professions
AI/ML Researchers & Engineers, UI/UX Designers, E-Commerce Platform Operators
article
Content Status
Full text indexed
hub
Related Papers
2 related papers