A1.13.2Pictorial depth cuesresearchdesign

Shading, size, and perspective provide relative depth

Aliases: cast shadow · relative size · familiar size · linear perspective · texture gradient

What it is

Beyond occlusion, the strongest and most stable depth cue, there is a whole family of weaker, more assumption-dependent monocular depth cues: cast shadows, relative or familiar size, linear perspective (parallel lines converging into the distance), and texture gradient (repeated texture units getting denser with distance). Collectively called pictorial depth cues, they share the property of working from a single eye and a single static image, but they deliver graded, relative depth information rather than occlusion's binary front-back call — and they matter most precisely when occlusion isn't available, i.e. when objects in a scene don't touch or overlap.

Why it happens

Each of these cues rests on a learned assumption about the world. Cast-shadow position and shape imply a light-source direction and how high an object sits above the ground or background, assuming a fixed-direction light source (the visual system has a default bias toward "light comes from above"). Relative or familiar size assumes that similar objects project a smaller retinal image as distance increases, which requires either a known real-world size or a comparable reference object. Linear perspective and texture gradient assume that, under standard projection geometry, parallel lines and repeated texture units converge and densify with distance. Because all of these cues rest on the visual system's built-in assumptions rather than a certain measurement, they degrade — or produce outright wrong depth judgments — whenever the underlying assumption is violated (unusual lighting direction, no size reference, an unconventional viewpoint). This is exactly how many depth illusions are constructed.

Studying it

  • Single-cue manipulation experiments: independently varying shadow direction, relative size ratio, or perspective convergence angle in an otherwise flat or ambiguous image, and measuring how judgments of depth or elevation shift accordingly.
  • Cue-conflict experiments: setting two pictorial cues against each other (e.g. shadow implying one height while size implies a different distance) and measuring which cue carries more weight in participants' judgments.
  • Typical independent variables: shadow direction and softness, the size ratio between compared objects, perspective convergence angle, texture density gradient.
  • Typical dependent variables: ranking of judged relative depth or height, perceptual judgment of 3D shape.

Where it stops holding

  • These cues deliver relative, comparative depth information, not the calibrated, precise measurement some other cues can provide — and whenever their underlying assumption is violated (unusual shadow direction, unknown true object size), the cue itself can produce a wrong depth judgment. This is exactly how depth illusions are built.
  • Without a comparison reference, these cues contribute nothing at all: a single isolated object against a flat, uniform-colored background — no visible shadow, no comparable object, no perspective lines or texture — carries no depth information from this group of cues.
  • This entry only covers the effect of these cues in isolation; how multiple pictorial cues combine or conflict when presented together is a separate question.

Applying it

  • Keep a single, consistent lighting-direction assumption across the interface: if drop shadows are used to express an element's elevation or layering, the global shadow direction should stay consistent — inconsistent shadow directions will make the interface "look off" even if users can't articulate why.
  • Provide a size reference, or keep object scale consistent, when size is used to convey layering or distance: relying on size alone to imply depth or importance, without a reference object or with a reference whose own scale keeps changing, prevents users from reliably reading the intended depth relationship.
  • When expressing pseudo-3D layouts (floating panels, perspective cards, AR previews), stack several pictorial cues together rather than relying on just one — each cue is individually weak and easily overridden, and combining cues substantially raises the odds that the intended layering is read correctly.
  • Verification: isolate a cue like shadow or perspective on its own (with stronger cues like occlusion removed) and test whether users can still correctly judge the intended stacking order or relative distance — confirming the cue is doing real work rather than being carried by occlusion.

Related

  • Same group: A1.13.1 Occlusion is the strongest and most stable depth cue · A1.13.3 Binocular disparity is only effective at close range · A1.13.4 Conflicting cues make depth judgments unstable
  • Nearby: A1.27 Perceptual constancy
  • Search terms: pictorial depth cue · cast shadow · relative size · linear perspective · texture gradient

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/A1.13.2