Low agreement coefficients diagnose ambiguity in the coding frame
Aliases: unreliable coding frame · low kappa · overlapping codes
What it is
A low agreement coefficient should not first be read as “one coder was careless.” When training and material difficulty are roughly comparable, a low rate is more often the coding frame’s ambiguity speaking: overlapping codes, two phenomena stuffed into one code, an unset unit of analysis, or a mass of material that only fits in “other.” Low agreement as codebook diagnosis treats the number as a signal about the grid, not as a personnel review.
Why it happens
If two competent coders repeatedly drop the same stretch into different cells, the cell boundaries do not work on this material. Typical sources: definitions that circle through synonyms; positive examples that are all prototypes and no negative examples; exclusive-code rules colliding with “multiple codes allowed”; different segmentation grain producing apparent category conflict. Rates are also warped by the number of categories and by base rates: one disagreement on a rare code wrecks the coefficient; a crude, very common code inflates it. A low number therefore points first to “this grid cannot stably cut the material,” and only second to execution.
Studying it
Do not stop at one overall coefficient. Emit confusion by code: which pairs steal the same stretches, which codes dump into “other.” Count unitizing disagreements separately from category disagreements. Run a definition-rewrite experiment on high-disagreement codes—split, merge, or add negative examples—and retest on new material; the frame is repaired only if the coefficient rises and the themes can still answer the research question. Recorded coder talk-alouds help separate “missed it” from “the rule allows both readings.”
Where it stops holding
Low coefficients sometimes do come from under-training, fatigue, or material so vague it has almost no shared referent (participants using conflicting words for the same thing). Repairing the manual will not fix that; the work has to go back to the field or the interview. Chasing a coefficient by forcibly merging codes that carry analytic tension yields a pretty number and empty themes. Different coefficients (percent agreement, kappa, Krippendorff’s α) treat chance agreement differently; comparing thresholds across studies is meaningless.
Related
- Same group: Q4.11.1 Agreement needs a shared codebook and explicit definitions · Q4.11.3 Reaching agreement can surface neglected categories · Q4.11.4 High agreement is not analytic value
- Adjacent: Q4.02 Coding and thematic analysis · Q1.14 Researcher bias and leading questions
- Search terms:
Cohen's kappa·Krippendorff's alpha·coding frame