Unexplained terms hurt more than long sentences
Aliases: jargon density · unexplained terms · lexical burden
What it is
How densely unexplained terms appear usually predicts outsider failure better than how long the sentences run. Eight words of jargon are harder than a thirty-word narration whose words already sit in the reader’s lexicon. Readability formulas reward splitting sentences. They do not reward filling holes in the proposition.
Why it happens
Each unknown term costs a lexicon lookup. A miss leaves a hole in the proposition. Syntactic working memory will often carry a native reader through a subordinate clause; a hole cannot be patched — the reader does not know what should have been there. So “undefined terms per hundred words,” relative to that reader’s lexicon, sits closer to the failure than mean sentence length.
Terms familiar inside the field are not load; they are compression. Density has to be relative to the lexicon. Absolute counts will call an expert page a disaster. The formula cannot see this: it sees syllables. Split the long sentence, leave the terms, and the grade improves while every outsider hole remains. Leave the sentence long, ground each term at first mention, and comprehension can pass. The comparison is which predictor carries the larger main effect. It is not a lesson in how to write sentences.
Studying it
Factorial: hold sentence length, vary the count of undefined terms; then hold terms, lengthen sentences. Expect a larger main effect of terms, and a smaller one for length except at extreme embedding.
Independent variables: unexplained terms per hundred words, mean sentence length, whether readers are inside the field. Dependent variables: comprehension items, cloze, self-report of which word blocked them.
Code “unexplained” tightly: no on-page definition, no link to one, not uniquely recoverable from context. A term already grounded above does not count.
Where it stops holding
When the readers are the field’s experts, high density is efficiency; lowering it unpacks chunks they already hold. Verse, brand names and required statutory names are not a terminology problem. Stacked negation and deep embedding can crash understanding on their own; structure is then a main effect, not “always the terms.” Machine-smoothed short sentences whose terms are simply wrong look low-density and still fail — that is error, not density, and a different correction.
Applying it
- On public-facing pages, count unexplained terms. Do not only watch mean sentence length or a grade.
- Load-bearing terms must be grounded by something on that page at first appearance (a definition, an example, a link). Do not use a sentence split as a substitute for grounding.
- How to check: move sentence boundaries until the formula passes, but do not ground the terms. Outsiders’ paraphrases still fail — the wrong variable was changed. Then ground the terms and leave the sentence boundaries. If paraphrase now passes, density was the variable.