L6.07.2post-hoc recommendation hallucinationdesignresearch

Reasons must come from the actual ranking basis; a post-hoc story is another kind of hallucination

Aliases: rationalised why · post-hoc explanation · invented because

What it is

The ranker scores first. Then a separate language model or copy template, looking at cover and title, invents a plausible “because.” That “because” never entered scoring. Post-hoc recommendation hallucination is a reason generated as a story after the result, decoupled from the basis, still wearing the form of an explanation.

Faithfulness is “this sentence points at a real contribution.” Here there is a further cut: even the production order is effect then cause. Hallucination is not only a chat-model event; it lives in the reason rail.

Why it happens

Post-hoc generation optimises for sounding like speech, not for passing the scoring log. So it can cite unused features, invent acts the user never did, and dress an ad slot as taste. People simulate the system from that story; the next forecast skews. Worse, the story is often coherent, easier to believe than a true but ugly basis (quota, paid placement).

When the generation pipe and the ranking pipe are cut apart, rebuttal also fails: the user denies the “because” in the story, and the actual scoring features do not move. Rebuttability needs reason and basis on one rope. Post-hoc invention cuts the rope and leaves decoration.

Studying it

Contrast three reason sources: contribution pulled from the ranking log, a template filled with real features, a generative model writing “because” from metadata only. Code: mentions of real contribution, mentions of unused features, invented behaviour. Dependent variables: whether people can forecast the next item from the reason, whether ranking changes after the reason is denied, rate of judging the reason “made up.”

Do not use readability as the endpoint. Post-hoc copy will almost always read better. Report alignment with the log, and whether denying the copy changes treatment. Faithfulness research asks whether the content is true; here the question is whether production ran effect-before-cause.

Where it stops holding

A human-curated list with no scoring log should be labelled “editors’ picks,” not dressed as a personal basis. When features cannot be disclosed, it is allowed to say “other unlisted factors” — that admits a gap; it is not inventing a substitute story. This entry targets the post-hoc generation pipe. It does not treat semantically empty boilerplate, and it does not treat reasons as selling.

Applying it

  • Allow reasons only from features or neighbours used in this scoring pass. If a generative model is in use, its input must be that feature list, not cover and title alone.
  • Sample weekly: align reasons with scoring logs; take down sentences that hit unused features or invent behaviour.
  • Check: deny the on-screen “because”; ranking should move. If the story is denied and rank does not move, that is post-hoc hallucination.

Related

  • Same group: L6.07.1 A reason helps the user decide whether an item is worth opening, not to persuade them to open it · L6.07.3 Reasons expose data sources and can make users feel tracked · L6.07.4 Reasons make recommendations rebuttable, so users can correct a wrong profile · L6.07.5 The same sentence on every item conveys no information
  • Nearby: L6.01 Recommendation Reasons · L3.03 Hallucination and the Fact-Checking Burden · L5.08 Counterfactual Explanations
  • Search terms: post-hoc recommendation hallucination · rationalised why · faithful explanation pipeline

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/L6.07.2