A rating’s discrete grain sets how much distinction it can express
Aliases: scale points · star rating · thumbs versus stars
What it is
A rating control packs attitude or quality into a finite set of discrete bins. Granularity is how many bins: two (up/down), five stars, ten, a hundred. The count sets how much distinction a person can express, and how much distinction later analysis can claim. Too few bins crush different feelings into one cell; too many make the gap between neighbours smaller than a person can reproduce, and the score starts to jitter. Finer stars are not more professional; they are a step chosen between expressiveness and repeatability.
Half-stars turn five bins into ten while still looking like five stars; the grain has already changed.
Why it happens
Absolute judgement stably carves only a limited number of categories. Beyond about seven untrained categories, subjective distance between neighbours overlaps, and the same person will score differently a day later. Too few bins do the opposite: the middle of three swallows weak-positive and weak-negative, and the product cannot tell “fine” from “barely usable.” Stars also carry cultural calibration — five often reads as a perfect score, four as mild complaint — so grain rides on that anchor rather than on a linear ruler.
Displaying a mean compresses grain again: internally a ten, externally five stars plus one decimal. The distinction people typed and the distinction others read are not the same ruler. If input is five bins and display is 4.73, readers infer precision the rater never gave.
Studying it
The classic move is to vary the number of scale points and watch test–retest reliability, neighbour-bin use, and whether extremes are abandoned. The task can be scoring the same objects twice, days apart.
Independent variables: number of bins (2 / 5 / 7 / 10), half-steps allowed, presence of verbal anchors. Dependent variables: retest agreement, occupancy of middle versus extreme bins, time to finish, whether people can later put the difference between two neighbours into words.
Do not look only at variance of the mean. Large variance may be real difference or jitter from too many bins. Split out participants who cannot name the gap between neighbours; that is closer to a grain failure than the group standard deviation.
Where it stops holding
Experts in their domain (wine, clinical scales) can use more bins because the categories were trained; handing the same ten-point scale to a one-time visitor yields noise. Users whose culture has no “four-star” habit treat five stars as a switch (good/bad), and grain exists only in name. If “unrated” must be allowed, zero stars and one star cannot share one empty glyph, or unrated leaks into the mean. Comparison tasks (which of these two is better) are more stable than absolute scores; if what you need is an order, do not harvest it indirectly from five stars.
Applying it
- Ask how many bins later decisions need (ship / improve / kill), then collect with no fewer and no more bins than that.
- Turn on half-stars only when analysis truly needs ten bins; still calling it five stars misleads the rater.
- Decimal places on a displayed mean must not exceed the precision the input grain allows.
- How to check: have the same person rate the same object a day later. Neighbour-bin bouncing without a named difference means fewer bins; a pile-up in the middle while talk is actually valenced means more bins or better anchors.