A long-term holdout that never receives the change is a common way to observe long-term effects
Aliases: holdout · never-shipped control · long-run baseline
What it is
To see a change’s net effect months later, a common practice is to keep a long-term holdout that never receives the change: after full ship they still walk the old path until the observation window ends. If the short experiment ends and everyone is shipped, the contrast disappears, and there is no longer a “had we not changed” in the later world. A holdout is not running the experiment a few extra days. It is giving up letting some people enjoy (or suffer) the new design, in exchange for a baseline that is still clean.
Why it happens
Full ship kills the counterfactual. A long-term effect has to compare, on the same calendar, people who changed and people who did not. Season, spend, and competitor moves on that calendar should hit both sides, so the difference can be attributed to the change. If the control is merged in week two, a trust drop starting in week three can only be compared with history, and history is no longer the same world. The costs of a holdout are opportunity, fairness, and engineering: the new feature must be switchable per person, and someone has to explain why some people never see it. It can observe long-term effects precisely because it refuses the shipping instinct that “everyone should end up the same.” The sample also has to survive long-window attrition, or the control collapses on the due date.
Studying it
Pre-compute remaining holdout sample and minimum detectable effect at the due date, and size the holdout fraction on post-attrition n, not day-one n. Keep the primary analysis intention-to-treat: mid-flight activity changes do not move people out of assignment. Record contamination: did holdout users meet the new design through sharing, support, or media. At the end, compare holdout and treatment on pre-registered long-term endpoints, and report external events during the holdout that should have hit both sides. Qualitatively follow up holdout users to confirm they were still on the old path.
Where it stops holding
Safety fixes, legal mandates, and changes that must be identical for everyone cannot hold out. A holdout that lasts too long becomes discrimination against those users, especially when the new design is clearly better; there should be an ethical cap and compensation. When several changes stack, one global holdout cannot tell which change produced the long-term effect; stratify by change or accept that only the bundle can be estimated. A holdout also cannot observe the equilibrium after the whole network has changed, such as the new industry normal after everyone receives noisier notifications.
Applying it
- For changes that may have effects at month scale or longer, write holdout fraction, due date, and long-term endpoint into the ship plan; do not merge the control before the due date.
- Guarantee per-user off switches, and periodically sample whether holdout users are still on the old path.
- Write the holdout’s opportunity cost into the decision; do not shrink the control quietly under full-ship pressure.
- Until the due analysis is done, do not announce long-term success with mixed totals after a short-window full ship.
Related
- Same group: Q6.10.1 Organizational evaluation cycles are usually shorter than the time UX harm takes to appear · Q6.10.3 Decision-makers systematically underestimate uncertainty in long-term effects · Q6.10.4 Short-term metric gains bought with manipulative design erode long-term trust
- Adjacent: Q6.06 Long term and short term · Q3.04 A/B testing
- Search terms:
long-term holdout·holdout group·long-run experiment