Breaking self-reinforcement needs a signal from outside the loop
Aliases: exogenous exploration · editorial injection · off-loop signal
What it is
Every step inside the closed loop confirms the last one. To walk out, something has to arrive that was not produced by this user, this display, this click: an external loop-breaking signal — an editorial list, popularity from other cohorts, a deliberate exploration slot, a correction the user typed, a random or structured injection independent of the current model. Without an outside term, self-reinforcement has no breakpoint.
“Train another round” is not an external signal. The new round still eats labels from the same loop.
Why it happens
The fixed point of self-reinforcement is: the display policy produces behaviour, and behaviour trains almost the same policy. The point can be so stable it looks like taste; it is only unperturbed. Breaking it needs an input orthogonal to current parameters, or the gradient keeps walking the old direction.
External is not the same as arbitrary. Editorial slots bring someone else’s judgment; cross-cohort popularity brings someone else’s clicks, not this person’s echo; exploration slots bring a query the product designed. What they share: the label-generating process does not sit inside this user’s loop. Mix those labels with in-loop labels in one loss and leave them unmarked, and the next round swallows the external signal — so it has to be protected as its own channel.
The immediate experiential cost of exploration, and offline evaluation’s bias toward the logging policy, are finer operational issues. Here the requirement is only: the breakpoint must come from outside the loop.
Studying it
On users already in a steady state, inject an external source: an editorial rail, cross-group head titles, a small share of model-independent items, against “keep the loop closed.” Watch whether the steady list moves, whether users read the injection as error, and whether the move rebounds when injection stops. Independent variables: type of external source, injection share, whether it is weighted separately in training. Dependent variables: drift of the list from the steady state, rebound time, subjective “I was interrupted.”
Regret analysis from exploration–exploitation can describe how efficiently you inject, but minimum regret is not automatically the best experience. The HCI question for an external signal is whether people can understand “why did this appear,” and whether the loop snaps shut the moment injection stops.
Where it stops holding
When the user has already supplied a strong outside signal through filters, search, or subscriptions, more injection stacks. Compliance and safety lists are another kind of external signal, aimed at bounding the loop rather than opening it. When the item pool is tiny and the external source overlaps the inside, injection will not show. This entry argues that the breakpoint has to enter from outside the closed loop. It does not treat how accidental clicks freeze into a profile, and it does not treat offline evaluation eating its own logs.
Applying it
- Give the external source a channel retraining cannot swallow: its own rail, or a separately weighted, capped term in the loss.
- Label the source: “editors are pushing this,” “other regions are watching,” so the foreign item is intelligible rather than dressed as “we know you better.”
- Check: stop all external injection for two weeks on a steady cohort, then open an editorial rail for a week. If the list pins during the stop and only moves when the rail opens, the earlier stability came from missing a breakpoint, not from having found the taste.
Related
- Same group: L6.04.1 What is shown shapes behaviour, which then reshapes what is shown · L6.04.2 The loop amplifies whatever bias arrived first
- Nearby: L6.09 Feedback Loops and Preference Entrenchment · L6.08 Filter Bubbles and Diversity · L6.13 Negative Feedback Channels for Recommendations
- Search terms:
external loop-breaking signal·exogenous exploration·editorial injection