Deception at Scale: Deceptive Designs in 1K LLM-Generated E-Commerce Components
Authors
Paper Title
Deception at Scale: Deceptive Designs in 1K LLM-Generated E-Commerce Components
Publication Info
- Topic area: Large Language Models (LLMs) and deceptive design patterns in e-commerce interfaces.
- Keywords: LLMs, deceptive designs, dark patterns, e-commerce, interface interference, user-centered design, business incentives, human values, prompt engineering, front-end code.
Background and Problem
- Problem / challenge: LLM-generated front-end code often embeds deceptive designs, potentially scaling harmful practices unintentionally.
- Significance: Deceptive designs can harm users financially, invade privacy, and create cognitive burdens, and their automated generation risks widespread deployment without oversight.
- Motivation and related work: Prior studies have documented deceptive designs in human-designed interfaces and found initial evidence of dark patterns in LLM-generated code. However, systematic, large-scale investigations into the frequency, strategies, and mitigation of deceptive designs in LLM outputs were lacking.
Solution
- Proposed approach: A large-scale audit of LLM-generated e-commerce components to identify deceptive designs, assess influencing factors, and test strategies to reduce their frequency.
- Novelty:
- First large-scale audit of deceptive designs in LLM-generated front-end code.
- Introduction of design rationales to infer intent and interaction assumptions.
- Testing of stakeholder interest prompts and user-centered strategies to mitigate deceptive designs.
- Development of a dataset of 1,296 annotated e-commerce components.
- Procedure and key techniques:
- Study 1: Generated 1,080 components across 15 e-commerce types using four LLMs under three stakeholder interest conditions (business, user, baseline). Annotated designs using Gray et al.’s taxonomy.
- Study 2: Tested three new user-centered prompt strategies (direct mitigation, deception definition, human values) on six components using two LLMs. Compared outputs for reductions in deceptive designs.
Results
- Concrete findings:
- 55.8% of components contained at least one deceptive design, and 30.6% contained two or more.
- Interface interference was the dominant strategy (63.2% of instances), followed by forced action (15.3%) and social engineering (14.0%).
- Business interest prompts increased deceptive designs by 15.8 percentage points, while user interest prompts reduced them by 5.8 percentage points.
- Human value prompts reduced deceptive designs by 15.3 percentage points compared to baseline.
- Advantage over baselines:
- DeepSeek-V3 produced significantly fewer deceptive designs (46.3%) compared to Gemini 2.5 Pro (60.7%), Grok 3 Beta (60.4%), and GPT-4.1 (55.9%).
- Human value prompts were the most effective mitigation strategy, outperforming direct mitigation and deception definition.
- Experiments / evaluation:
- Study 1: Factorial design testing 15 components, 4 LLMs, and 3 stakeholder interest conditions. Manual annotation of 1,080 outputs.
- Study 2: Stratified sampling of 6 components and 2 models under 3 new user-centered strategies. Manual annotation of 216 outputs.
- Limitations and future work:
- Focused only on e-commerce components; other domains may exhibit different patterns.
- Static components excluded JavaScript-based dynamic behaviors.
- Some annotation disagreements highlight the need for refined criteria and further study of controversial cases.
Summary
This paper presents the first large-scale audit of deceptive designs in LLM-generated e-commerce components, revealing that over half of outputs contained at least one dark pattern. The study identified interface interference as the dominant strategy and showed that business interest prompts significantly increased deceptive designs, while human value prompts were the most effective at reducing them. By introducing design rationales and testing mitigation strategies, the paper highlights systemic risks in LLM-assisted code generation and offers actionable recommendations for developers and providers to embed safeguards and human values into future systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
The Siren Song of LLMs: How Users Perceive and Respond to Dark Patterns in Large Language Models
CHI '26· Dark Patterns Recognition +2
- 67%
Deus Ex Machina and Personas from Large Language Models: Investigating the Composition of AI-Generated Persona Descriptions
CHI '24· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)