After a major redesign the old baseline is gone; a new comparison period has to be built
Aliases: baseline break · post-redesign control · new comparison window
What it is
A major redesign changes not a few controls but the path, habits, and reference points with which people finish work. Conversion, time, and satisfaction measured on the old path are no longer a historical baseline for the new one. Judging the new version needs a fresh comparison: hold some people on the old version, run a contemporaneous control, or let a stable period accumulate on the new version before later tweaks are judged. Stacking last year’s same-week number onto the new version compares two products, not the weather around one experience.
Why it happens
A baseline works only if, had nothing changed, the metric would have tracked the past. A major redesign breaks that counterfactual: information architecture, defaults, and visual hierarchy move together, and even the shape of errors is different. Novelty and learning further make the first stretch after launch not the new system’s steady state; comparing that stretch to the old steady state packs the difference with transients. Aligning on season (this year’s holiday versus last year’s) controls the calendar and does not rebuild the old interface. Every later small optimization that still cites the pre-redesign grand mean then mixes the redesign step with subsequent tweaks in one trend, both flattering and blaming the wrong change.
Studying it
Choose the control strategy before launch: concurrent holdback on the old version, staged exposure, or an explicit break and a new series. Write the break into the analysis and cut the trend chart there; do not join the line. Mark the first part of the new series as transient and date the start of the new steady state separately. Later experiments take that new steady state as their control, not the pre-redesign average. If the old version cannot be kept, a queue that never saw the new version or an almost-unchanged external task is a weak control, and the write-up must say the counterfactual is incomplete.
Where it stops holding
Local edits—one sentence of copy, one color—usually do not void the whole historical baseline, which can still be used if step definitions did not move. A forced full rollout with no control can be described, not given a causal reading. A new baseline should not be locked too early: locking while the transient is running writes a trough or a peak into “normal.” How long the comparison period lasts is still paced by usage, but the point here is that the object of comparison changed; stretching the window does not let you keep the old ruler.
Applying it
- Put “the historical trend breaks here” in the redesign notes, insert a break on the dashboard that day, and disable automatic year-on-year joining of old and new.
- Keep a concurrent old-version or delayed-full-rollout control until the new version meets a pre-defined steady state.
- For the first optimization experiment after the redesign, name the control as the new-version steady state, not last year’s number.
- Check: score the new version against the pre-redesign mean and against a contemporaneous control. If the two scores disagree, stop using the historical mean as a ship/no-ship call.
Related
- Same group: Q3.14.1 Early metric movement after a change can be transitory · Q3.14.2 Learning costs temporarily suppress metrics for existing users · Q3.14.3 Observation windows must be long enough · Q3.14.4 Novelty typically rises then falls; learning typically falls then rises · Q3.14.5 Segmented comparison is required to tell them apart · Q3.14.6 Engagement metrics pick up novelty more than task success
- Adjacent: Q3.18 Experimental design and controls · Q5.09 Staged rollout and pilots
- Search terms:
historical baseline·counterfactual break·post-redesign control
Cards in the same group
- Q3.14.1Metric movement right after a change may not last
- Q3.14.2Relearning costs for existing users pull metrics down for a while
- Q3.14.3The observation window has to outlast the transient
- Q3.14.4Novelty usually rises then falls; learning usually falls then rises
- Q3.14.5Telling novelty from learning takes a split by people; the blended curve confuses them
- Q3.14.6Engagement metrics absorb novelty more readily than task success does