LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
Authors
Research Background and Problem
-
What problems or challenges did the authors identify?
- Prompt engineering is a technique that uses natural language to guide large language models (LLMs) in generating specific responses. However, for general users, crafting effective prompts is often neither intuitive nor straightforward.
- Some services and platforms have emerged to provide optimized and recommended prompts, but these services pose potential risks: third parties can subtly manipulate the phrasing of prompts to influence the output of language models, embedding significant biases into the content.
- Previous studies have primarily focused on biases in language models related to data and usage processes, but there has been limited research on the subtle manipulation of prompts and its practical implications.
-
Why is this issue important?
- If prompt recommendation services are maliciously exploited, such as by subtly modifying prompts to manipulate model outputs, this could lead users to adopt specific biased viewpoints or choices, threatening their cognitive independence.
- Such attacks could be used for covert advertising, political propaganda, or spreading misinformation, with potentially severe societal, economic, and political consequences.
-
Research motivation and related work
- While some research has examined the sensitivity of models to input variations and the biases in language models, few studies have explored the practical effects of attacks using subtle prompt modifications.
- This study aims to demonstrate that even seemingly innocuous prompt statements, when slightly adjusted, can guide language models to produce potentially biased outputs.
Solution
-
What methods or solutions did the authors propose?
- The researchers developed a weak perturbation method based on synonym replacement, substituting certain words in prompts with inconsequential but natural synonyms to achieve attack objectives, such as prompting the language model to mention a target concept (e.g., a specific brand or political party) more frequently.
- They proposed a loss function based on model output logits as an optimization metric to select prompts most likely to increase the appearance of target vocabulary in the generated output.
-
What is innovative about this solution?
- The study introduced a novel threat model that induces the generation of significantly biased outputs by merely altering the phrasing of prompts, while ensuring the prompts and responses remain imperceptible to users.
- The synonym database was improved to avoid generating semantically anomalous outputs due to inaccurate substitutions.
- Through user experiments, the study was the first to validate the stealthiness of such attacks and their influence on user preferences.
-
What are the implementation steps? What key technologies were used?
- Prompt Perturbation: Starting from unbiased prompts, a large number of new prompts were generated using synonym replacement.
- Attack Optimization: A target loss function was used to select prompts that maximized the probability of the target vocabulary appearing.
- Experimental Validation:
- Six open-source LLMs (e.g., Llama 2, Llama 3) were tested to generate responses containing the target concepts, and changes in the generation probabilities of the corresponding vocabulary were measured.
- User experiments were designed to validate the practical effectiveness of the attack, including evaluating the stealthiness of the prompts and responses and their influence on user preferences.
Research Outcomes
-
What specific results were achieved?
- Across all tested models, synonym replacement achieved up to a 78.3% increase in generation likelihood, significantly raising the probability of target words appearing.
- User experiments showed that the observed prompts and responses were statistically indistinguishable from unperturbed prompts and responses, but users were more likely to notice the target brand when exposed to modified prompts.
- The attack design not only successfully induced bias but also, in many cases, led users to prefer the modified prompts and responses.
-
What advantages does it have compared to existing solutions?
- Unlike attack methods that require direct access to model weights, this method operates entirely in a black-box scenario, requiring no knowledge of the internal mechanisms of language models, making it highly practical.
- The study validated the effectiveness and danger of this low-cost, highly stealthy, and easily implementable attack method.
-
What were the experimental or evaluation results?
- For users, the vast majority of modified prompts and responses exhibited equivalent semantic clarity and usability compared to unmodified ones.
- Assuming users trust third-party suggested prompts, the attack successfully increased the model’s recommendation probability for the target brand or concept while reducing the frequency of other concepts (if desired).
-
Limitations and future directions
- Limitations:
- The study primarily focused on open-source LLMs, such as Llama-2, rather than closed platforms like OpenAI’s ChatGPT.
- User experiments were conducted in specific scenarios (e.g., shopping-related prompts), while other contexts (e.g., political topics) remain unverified.
- Future directions:
- Explore more complex attack scenarios and threat models, as well as implementation in other languages or non-English environments.
- Investigate the transferability of attacks across language models (cross-model transferability).
- Develop dedicated defense mechanisms and long-term auditing strategies to detect and mitigate such prompt-based attacks.
- Limitations:
This study significantly contributes to understanding the relationship between prompt design and language bias, while raising serious concerns about prompt security. It highlights the need for vigilance and the exploration of countermeasures in real-world applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can slightly modified prompts induce large language models to generate significantly biased outputs?Category: Bias and Fairness in Large Language Models and GenAISimilar questionsarrow_forward
- What methods can verify the hidden influence of prompt modifications on user preferences and generated content?Category: Bias and Fairness in Large Language Models and GenAISimilar questionsarrow_forward
- What are the applicability and limitations of prompt-modification attacks across different language models and scenarios?Category: Bias and Fairness in Large Language Models and GenAISimilar questionsarrow_forward
Practical Problems
1- Lay users struggle to judge whether prompt recommendation services embed hidden bias.Category: Bias and Fairness in Large Language Models and GenAISimilar questionsarrow_forward
- 83%
Prestige and Prejudice: How the Interplay of Recruiting Work and Algorithms Reinforces Social Inequities in Software Engineering
CHI '25· AI Ethics, Fairness & Accountability +1
- 71%
The Siren Song of LLMs: How Users Perceive and Respond to Dark Patterns in Large Language Models
CHI '26· Dark Patterns Recognition +2
- 67%
A Canary in the AI Coal Mine: American Jews May Be Disproportionately Harmed by Intellectual Property Dispossession in Large Language Model Training
CHI '24· AI Ethics, Fairness & Accountability +1
- 67%
Deus Ex Machina and Personas from Large Language Models: Investigating the Composition of AI-Generated Persona Descriptions
CHI '24· Human-LLM Collaboration +1
- 67%
AI Mismatches: Identifying Potential Algorithmic Harms Before AI Development
CHI '25· AI Ethics, Fairness & Accountability +1
- 63%
"Please, don’t kill the only model that still feels human": Understanding the #Keep4o Backlash
CHI '26· Generative AI (Text, Image, Music, Video) +3
Based on Jaccard similarity of research subtopics & professions (≥60%)