Item wording shifts the response distribution
Aliases: wording effect · loaded stem · distribution shift from phrasing
What it is
In a questionnaire, question wording is not a neutral rendering of a concept. It specifies the object, evaluative direction, and acceptable intensity the respondent uses to construct an answer. The same attitude in the same people will move its agreement rate and mean when “restrict” becomes “ban,” or “spend” becomes “waste.” The distribution can shift entirely from measurement, not from the population or the product. Reading a results table without the stem treats a wording effect as a factual difference.
Why it happens
Most answers are not retrieved scores. They are assembled on the spot from the stem, examples, and implied norms. Evaluative verbs in the stem nominate what ought to be opposed. Vague frequency words (often, sometimes) map to different cutoffs for different people. Examples shrink the attitude object to the subclass in the example. Social desirability also lives in word choice: “Do you carefully read the terms?” pushes answers toward a respectable pole more than “How long did you stay on the terms page?” Wording therefore moves the answering process, not a pre-stored true score. The distribution follows the process.
Studying it
Use a split-ballot wording experiment: randomly assign two items that differ only in a keyword or example within the same population, and compare choice shares, means, and “not sure” rates. Record the swapped word, the added example, and evaluative direction. The dependent variable is location and spread of the distribution, not which version is “correct.” Cognitive interviews ask people to restate the item and explain the chosen option, locating whether the object, the direction, or a threshold word was misread. Shifts found in pretest belong in the study’s limits; they should not be hidden by picking the smoother stem.
Where it stops holding
Expert samples hold terminology more stably, so wording effects shrink, and that stability does not transfer to end users. Factual behavior items (did you open X yesterday) resist evaluative verbs better than attitude items, and still move with time windows and examples. Translation introduces a second wording layer, which is out of scope here; even inside one language, dialect and jargon can shift a distribution. That an effect exists does not make it a legitimate lever: a distribution produced by extreme loaded words has no decision value.
Applying it
- Replace every abstract term with an observable object and a time window, then delete evaluative verbs and check whether the options still make sense.
- Pretest two versions of a key item that differ by one word; report both distributions. If the shift exceeds a pre-set threshold, rewrite or report both; do not pick the prettier version for the main survey.
- Replace frequency words with numbers or ranges (times per week, last seven days) so “often” is not mapped differently by each person.
- Paste the stem into the results table header before analysis; an indicator whose attitude object cannot be read from the stem does not enter an external claim.