Within-subjects designs need order effects, and usually counterbalancing
Aliases: order effect · counterbalancing · within-subjects carryover
What it is
The same person meeting several variants in sequence is a within-subjects design. Scores from the second meeting onward mix practice, fatigue, contrast, and residue: the interface already used changes how the next one is read. That systematic bias from sequence is an order effect. Counterbalancing assigns different orders to different people (or uses a Latin square or similar), so order is cancelled across variants as far as the structure allows, rather than assuming everyone will “just adapt.” A shipped product whose metrics rise then fall is a different time process, not the order of conditions inside one session.
Why it happens
Practice makes later tasks faster and cleaner; fatigue and boredom make them worse; contrast makes the second variant be judged relative to the first rather than to a true default. Irreversible residue is harder: once search is learned at the top right, failure on a top-left variant is no longer an unbiased measure of top-left. Shuffling order balances it in expectation; in a small sample one order can still pile into one arm. If order is not recorded and not estimated, the within-person gap ties variant to position. Counterbalancing does not delete order; it stops order’s main effect from impersonating the variant. A variant-by-order interaction (B looks good only after A) will be masked by cancellation and needs to be reported on its own.
Studying it
Fix the order structure in advance: full counterbalance, Latin square, or random order, and say whether each order cell will fill at the planned n. The analysis estimates variant, position (which meeting), and variant × position when needed. Critical tasks can add a washout or an unrelated filler, then check whether the later condition still carries the earlier condition’s error pattern. Report n per order and the variant gap inside each order; if one order drives the pooled call, downgrade the main claim. The power advantage of within- over between-subjects is conditional on order being absorbed by design, not a free gift.
Where it stops holding
Some treatments cannot be unlearned (a price, a spoiler, a shortcut). Counterbalancing cannot return people to ignorance; switch to between-subjects or use first exposure only. When order is the question (onboarding sequence, tutorial order), do not cancel it—manipulate it. When a live product cannot wash out memory of the previous version, lab within-subjects conclusions need the residue risk written down. Two very short trials sometimes carry little order; still record order so it can be tested.
Applying it
- Write the order table into the within-subjects protocol; do not send everyone through the same sequence (always old, then new).
- Report “which meeting” alongside variant; a draft with only variant means and no position check goes back.
- Default irreversible information (answers, prices, layout memory) to between-subjects, or analyze first contact only.
- Check: rerun on the first order only. If that call contradicts the counterbalanced result, do not write an unbalanced order gap as a variant win.
Related
- Same group: Q3.18.1 Without a control, treatment and time trend cannot be separated · Q3.18.2 Randomize at the unit where the effect occurs · Q3.18.4 A group gap is a treatment effect only after other variables are controlled
- Adjacent: Q3.04 A/B testing · Q3.14 Novelty and learning effects
- Search terms:
order effect·counterbalancing·carryover