High agreement is not the same as analytic value
Aliases: reliability is not validity · high-kappa trap · stable but shallow codes
What it is
High agreement among coders shows only that this grid can be reused stably on this material. It does not show that the cuts fall on the research question, nor that the resulting distribution is worth putting in a conclusion. Agreement is not analytic value: one can reliably tag “mentioned the interface” and still get nowhere on “how people understand risk.” Treating a high coefficient as proof of theme quality mistakes a repeatable operation for a warranted interpretation.
Why it happens
The easiest codes to agree on are often the shallowest: whether a word appeared, whether a step was completed, whether affect was positive or negative. Shallow codes reduce judgment, the coefficient rises, and information falls. Merging distinctions that carried tension, or dropping theoretically crucial codes that are hard to segment, also inflates the number. Agreement answers “would another person trained the same way mark it like this.” Analytic value answers “does marking it this way separate rival explanations and connect back to the question.” The two can travel together; they can also come apart completely.
Studying it
Beyond reporting a coefficient, require every theme slated for use to answer three questions: which phenomena it now keeps apart that could have been mixed, which raw stretches it depends on, and which part of the research question becomes unanswerable if it is dropped. Label high-agreement codes that cannot answer those three as descriptive indexes, not themes. Keep a difficult code as a foil: if its coefficient is lower but it unlocks a key mechanism, repair the definition rather than delete it for the number. Ask a colleague who did not code to look only at theme names and definitions and to imagine a counterexample; a high-agreement code that admits no counterexample usually has no analytic edge.
Where it stops holding
Without some repeatability, analytic value cannot be checked by anyone else, so coefficients at chance remain a problem. What is denied here is the substitution “high therefore good,” not agreement work as such. Some evaluative or compliance coding aims at stable classification; high agreement is then part of the value—and one still has to show the classification helps a decision. Thematic analysis that publicly declines to care whether anyone else can follow has already left contestable research.
Related
- Same group: Q4.11.1 Agreement needs a shared codebook and explicit definitions · Q4.11.2 Low agreement coefficients diagnose ambiguity in the coding frame · Q4.11.3 Reaching agreement can surface neglected categories
- Adjacent: Q4.02 Coding and thematic analysis · Q4.10 Transferability of conclusions
- Search terms:
reliability versus validity·analytic value·intercoder agreement