Q4.02.1audit trail from codes to datadesignresearch

Every code must lead back to a source extract

Aliases: traceable coding · anchored excerpts · unsourced codes

What it is

A report that says “users feel lost” with no timestamp, quote, or session is a précis, not a coding result. An audit trail from codes to data means each code can be walked back to the stretch that produced it: which record, which span, under what conditions. A code indexes a fragment; it does not replace the fragment as a new fact. Break the chain and later themes, personas, and opportunity lists spin without material.

Why it happens

Coding compresses. Compression drops speaker, condition, negation, and hesitation; unless a pointer stays on the code, the compression is irreversible. Teams then pass code names around as if the names were observations. When someone asks who said what, in which session, there is nowhere to go, and the claim is held up only by social agreement. A tidy codebook in software is fake tidiness if the file stores names without anchored excerpts. Traceability makes negation possible: produce the fragment, and the code can be changed, split, or retired.

Studying it

Sample backward from the final code list: for each high-use code, draw several items and check that a locatable transcript or video span exists and that the name and the span still co-refer. Compare an archive of frequencies only with an archive of code + excerpt + source, and see whether a reviewer can independently audit. Count unlinkable codes as analytic loss. Outcomes: locatability rate, recode rate after walking back, and the share of unsourced claims in the write-up.

Where it stops holding

Field jottings may temporarily lack a full citation, but they must be anchored before they enter the formal code list or they remain prompts, not codes. When law or ethics requires destroying identifiable source text, the trail shrinks to de-identified excerpts and situational description—it does not shrink to “we remember.” A thematic write-up may synthesize many fragments, but the synthesis should still point to a listable set of sources rather than a floating tag. Machine-assisted coding that emits label distributions without span alignment is equally unreviewable.

Applying it

  • Store at least four fields per coded instance: source file, location (time or paragraph), excerpt, code name.
  • Do not put a code cloud on a slide without a playable excerpt beside any claim that matters.
  • In review, pick a code at random and open the source span on the spot; if it will not open or does not match, the code does not enter the conclusion.
  • When changing tools or handing off, export the anchored excerpt table before exporting frequencies.

Related

  • Same group: Q4.02.2 Multiple coders require an agreement check · Q4.02.3 A theme is an interpretation, not a tally
  • Adjacent: Q4.01 Affinity diagramming · Q1.07.3 Record wording before interpretation
  • Search terms: audit trail from codes to data · traceability · qualitative coding

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/Q4.02.1