A short-horizon lift can be long-horizon harm
Aliases: short-term metric · long-horizon harm · experiment window mismatch
What it is
The primary metric in an A/B test almost always sits in a short window: clicks, conversions, or same-session completion over days to weeks. A lift inside that window can coexist with harm outside it—returns, complaints, next-month retention dropping, trust drawn down. When the experiment declares a winner, the long-horizon outcomes have not happened yet, so the harm cannot enter the decision. The problem is a mismatch between metric time and effect time, not a failure of the random split. Writing seven-day conversion as “a better experience” ends the evaluation on a photograph that is still being taken.
Why it happens
A short window selects behaviors that can change immediately: a more urgent button, a fuller default tick, an earlier permission prompt. Those moves pull people who would later refuse into a success event now; success enters the numerator, refusal is pushed past the window. Harm has a longer time constant: billing cycles, habit, word of mouth, regulatory complaints. The closer the primary metric is to the top of the funnel, and the shorter the window, the easier this temporal arbitrage is to win. Shipping to 100% then cuts the control, so even if long-horizon harm appears, the other arm that would have made it attributable is gone, and argument falls back on post-hoc stories.
Studying it
Pair the short-window primary with long-window guardrails in advance (seven-day conversion against thirty-day retention, refunds, repeat complaints). At the short window the experiment may only go “pending”; it declares after the guardrail date. Analysis should report both sets, not let the short win overwrite the long. If the control arm cannot be held for the long window, keep a stratified long follow-up and admit the drop in causal strength. Replay of historical experiments can show which arm a short window would pick versus which arm the long window would pick.
Where it stops holding
For one-shot tasks with no onward relationship (an event page, a single lookup), long-horizon harm may not exist and a short window is enough. Subscription, credit, health, and recommendations—domains with delayed cost—carry mismatch as a default risk. Longer is not automatically better: a window long enough that the business has changed three times empties the control arm of meaning. Guardrails must be named before the test; fishing in the long window for a drop after the date is a different kind of selection.
Applying it
- Freeze in the protocol: a short-window metric may only trigger “keep watching.” Full ship waits until guardrails mature without deterioration.
- For changes to default commitments, price display, or permissions, the guardrail must cover at least one full billing or use cycle.
- A variant that has won the short window but not yet reached the long window does not enter the quarterly win list; it stays pending.
- Replay the last two full ships. If the short metric rose and retention or complaints later worsened, lengthen the default window for that class of change; do not answer with an even shorter test.
Related
- Same group: Q3.04.1 Random assignment is what licenses a causal reading · Q3.04.2 A/B tests compare built variants; they do not invent new ones · Q3.04.3 Local A/B optima can hide structural problems
- Adjacent: Q3.12 Funnel and retention analysis · Q3.11 Logs and instrumentation
- Search terms:
metric horizon·guardrail metric·short-term lift