Long-term effects require long-horizon experiments
Aliases: long-running experiment · delayed-effect window · long-horizon A/B
What it is
Habit formation, trust, learning curves, and novelty decay do not finish inside a few days of contrast. To see those long-term effects, the experiment’s duration must cover the effect’s own time scale, not the iteration meeting’s time scale. That is a long-horizon experiment. A two-week conversion contrast can answer “will they tap more right now.” It cannot answer “will they still leave notifications on in three months.” Reading a short experiment’s null as “no long-term effect” treats the observation window as the phenomenon.
Why it happens
Effect curves have shape. Novelty lifts use in the first days and then falls; learning shows up as speed only after the first failures; trust moves after several keep-or-break events. A short experiment cuts the start of the curve, and the start can even have the opposite sign of the later part. Some effects also need calendar cycles: a billing date, a term, a season, an annual fee. If the experiment ends before those cycles, the corresponding behavior has not happened, and the variance is full of not-yet-due blanks. A long horizon is not the same daily metric drawn for more days; it is giving people in the denominator a chance to live through the events that change behavior. The costs are contamination, attrition, and mid-flight redesign, so a long-horizon experiment has to decide in advance who stays unchanged and what is allowed to happen in between.
Studying it
Write the minimum observable window for the claimed long-term effect first (two billing cycles, one full novice-to-fluent path), then set duration; do not set two weeks and then ask what can be said. Analyze on cohort age rather than calendar day, so late joiners do not flatten the curve. Pre-declare which mid-flight redesigns abort the experiment and which can be absorbed by stratification. Power must be computed on the post-attrition sample at the long window; powering on week-one sample is systematically optimistic. Short-window results can feed an early-stopping rule; they cannot replace the due analysis.
Where it stops holding
Not every question needs a long horizon: copy errors, accessibility hits, and crash fixes can conclude in a short window. Lengthening every experiment blocks iteration. Some long-term effects cannot ethically be manufactured in contrast and need natural variation or retrospective designs. A long horizon still cannot speak to the world after the experiment ends; it only moves the window from too short to adequate. Mid-flight full-ship destroys the contrast, so long-horizon experiments demand more shipping discipline than short ones.
Applying it
- State the effect’s minimum observation window in the proposal; the experiment must not end before that window.
- Report long-term metrics on cohort age, mark cells that are not yet due, and do not conclude from not-yet-due data.
- During a long horizon, freeze redesigns that would contaminate the main effect, or register them as stratification factors.
- If the short window looks good and the long window has not arrived, the status is “not yet decidable,” not a half-announcement of “success, long-term pending.”
Related
- Same group: Q6.06.1 Short-term gains can be paid for with long-term harm · Q6.06.3 Mismatch between decision cycle and effect cycle is a common trap
- Adjacent: Q6.10 Long-term effects versus short-term metrics · Q3.04 A/B testing
- Search terms:
long-horizon experiment·long-running experiment·effect time scale