Once a metric is used for evaluation, teams optimize the metric rather than the goal it represents
Aliases: Campbell's law · performance gaming · evaluation-driven optimization
What it is
Once a metric enters performance review, promotion, or fights over resources, the rational move is to move that number, not the goal it claims to represent. That is incentive-driven metric gaming. Support closes tickets early to raise resolution rate; growth turns registration into a skippable default to raise conversion; research pops the survey at the moment of success to raise the score. The agents are often not villains; they are ordinary people trained by the scoring formula. Unlike the general claim about a measure changing status, the mechanism here is an incentive contract: the number converts directly into reward.
Why it happens
Evaluation turns a metric from an observer’s instrument into a scoreboard for the observed. Once the scoring rule is public, the action space is rewritten by the rule: paths that add points are searched densely; paths that do not add points are dropped, including paths that truly serve the goal but do not show in the score. Interfaces offer cheap points: change a default, change a denominator, exclude difficult users from the statistic, time the survey at an affective peak. The tighter a single-column score is bound to a person’s fate, the more precise the gaming. Goal language can still sit on a values poster, but the weekly meeting only asks about the score, so the poster cannot compete with the formula. Stronger incentives, clearer rules, and a larger action space detach metric from goal faster.
Studying it
Compare interface and process strategies actually taken on the same metric before and after it is written into individual or team evaluation, and code how many bypass the goal. Use mystery shoppers or a criterion outside the evaluation to estimate whether goal attainment rose in the evaluation period as well. Interviews should ask “how would you reliably raise the score,” not “how do you improve experience”; the first question produces the gaming menu. Watch revisions of the scoring rule too: when denominators or exclusion rules are rewritten in ways that favor the score, gaming has already settled into the institution.
Where it stops holding
No evaluation at all starves experience work of resources. The issue is not “do not evaluate”; it is not turning a single manipulable number into a personal voucher. Team-level evaluation with constraints is harder to game precisely than a personal single column, but it does not kill gaming. A statutory completion rate almost is the goal, so the gaming space is smaller; exclusion rules can still be worked. After gaming is exposed, adding punishment without changing the formula makes gaming more covert rather than returning work to the goal.
Applying it
- An experience metric written into personal performance must be paired with a criterion the same person does not control; a split between score and criterion voids that evaluation.
- Ban personal vouchers on numbers that can be lifted the same week by changing a denominator or a default.
- In the evaluation briefing, ask “how could we raise the score without improving the goal,” and list the answers as forbidden levers.
- When gaming is found, change the formula and the statistical scope first, then talk about attitude; do not only process the people who were caught.
Related
- Same group: Q6.09.2 A single proxy is easy to inflate without improving real experience · Q6.09.3 Using multiple proxies in combination reduces single-point gaming · Q6.09.4 Pair reverse metrics to monitor whether the proxy relationship has already distorted
- Adjacent: Q6.05 Metric manipulability · Q6.03 North star metrics
- Search terms:
metric gaming·Campbell's law·incentive distortion