As a target, NPS invites contaminating optimization
Aliases: NPS gaming · Goodhart effect · contaminated recommend score
What it is
When NPS moves from a descriptive number to a performance review, a bonus, or a ship gate, it obeys Goodhart's law: a measure that becomes a target ceases to be a good measure. Contamination comes from changing the measurement process, and need not come from a better product—coaching users that “only 9 and above count as recommend,” holding the survey back from people about to leave unhappy, appealing low scores away, bundling a gift with the rating. A rising score then measures that the organization learned how to produce high scores, not that recommend intent moved.
Why it happens
The recommend item is cheap to answer, the scoring rule is public, and once money and reputation attach, the local optimum is to operate on the score rather than on the experience. Front-line staff control when the survey is sent, to whom, and what is said beforehand—none of that sits in product code, and it decides the sample and how the stem is heard. Exclusion rules (“not a real user,” “malicious detractor”) executed by the people being scored systematically delete detractors. The product side can also raise recommend intent without raising use quality: a more visible “share with a friend,” a recommend step that becomes mandatory after task success. The tighter the stare, the wider the gap between measure and construct.
Studying it
Place the score trajectory beside a log of the measurement process: changes to send rules, exclusion rates, gift campaigns, dates when talk-tracks went live. If a jump aligns with a process change and not with behavioral indicators (completion, repurchase, complaints), prefer contamination as the explanation. Audit exclusion lists and agent scripts, and compute how far the score would move if dropped low scores were restored. An organizational contrast between teams whose pay depends on the score and teams whose pay does not can show differences in coverage and scripting. Construct-validity evidence is whether the score’s correlation with independent behavior (actual recommends, reuse) falls as incentive intensity rises.
Where it stops holding
With no pay link, frozen send rules, and exclusions run by an independent party, contamination pressure is lower; scripts should still be spot-checked. When a regulator or buyer demands to see NPS, the organization may be unable to drop the metric, but it can uncouple it from internal bonuses to lower the incentive to pollute. Complaint volume or churn used as a sole target will attract a different pollution; that problem is not unique to recommend scores. The claim here is the behavioral reaction after the score becomes a target, not sampling bias in when a survey is shown.
Applying it
- Keep NPS out of individual or front-line bonus formulas. If a process must be scored, score send-coverage and protocol compliance, not the number itself.
- Register changes to send rules, exclusion rules, and scripts as if they were experiments. Scores inside a change window do not enter the trend.
- Each month, restore excluded responses and recompute. If ranks flip, mark that month’s score unusable.
- When the score diverges persistently from complaints, churn, and task completion, freeze its decision rights and inspect the measurement process before the product.