J4.06.1readability formuladesignresearch

A readability score is a coarse filter, not a pass

Aliases: Flesch–Kincaid · SMOG · grade level · readability score

What it is

Readability formulas such as Flesch–Kincaid or SMOG estimate a grade from sentence length and word length (or syllables). They are a sieve, not a pass mark. A score that says “primary school” while outsiders still fail is the formula missing what actually binds: whether terms are grounded, whether reference is clear, whether the reader brings domain knowledge. Chinese has no syllable rule you can import; wrapping an English formula around it is coarser still.

Why it happens

The formulas assume long sentences plus long words equal hard. On English textbook corpora that correlation is enough to rank. Interface copy will game them: split a sentence in the middle and the grade drops while the proposition is untouched. They cannot see stacked negation, tables, icons, or a definition that appears only after a click. Difficulty in Chinese more often sits in leftover classical diction, abbreviations and unexplained names — none of which have a counterpart in English syllable counts.

Conformance that uses “lower secondary” as a top-level reading requirement is itself a coarse threshold: it admits a line exists; it does not admit that the line equals understanding. Treat the formula as a release gate and teams will split sentences to move the number. The monitor falls. The reader’s failure points stay.

Studying it

Three-way on the same text: formula score, cloze, comprehension with target readers. Report how weak the correlation is, and whether understanding moved after a split-for-score edit.

Independent variables: which formula, whether sentence boundaries were gamed, language (English / Chinese). Dependent variables: grade, cloze hits, task items; the result that matters is disagreement between score and understanding.

Formulas help at scale (the hardest 5% of a site). Whether a single page may ship is a paraphrase by the intended reader. Do not turn this card into a second style manual — wording is a different problem. This one only scores the metric.

Where it stops holding

Closed-vocabulary technical fields inflate the score because terms are long; experts do not find them hard. Tiny UI strings are too short for a stable grade; the number is almost meaningless. Literature, humour and irony are outside the training genre. Automation that reports a score without saying which comprehension test it ever correlated with sells a sieve as a ruler.

Applying it

  • Use formulas as site-wide monitors. A red line means “go look,” not “edit until the number passes.”
  • Before release, have target readers paraphrase the load-bearing propositions. If the score dropped and the paraphrase still fails, the metric moved, not the difficulty.
  • How to check: take a piece the formula calls easy whose terms are unexplained, and give it to someone outside the field. If they cannot read it, the score was never a verdict.

Related

  • Same group: J4.06.2 Unexplained terms hurt more than long sentences · J4.06.3 Who is reading sets how hard the text may be
  • Nearby: J4.05 Plain Language · A11.05 Domain knowledge and terminology comprehension
  • Search terms: readability formula · Flesch–Kincaid · grade level

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/J4.06.1