Collaboration requires confirmation that “we are looking at the same place”
Aliases: shared attention · joint visual attention · shared visual focus
What it is
Joint attention is more than two people independently seeing the same object: each has evidence that the other is attending to it. “We are looking at the same place” establishes a shared starting point for pointing, explanation, and modification. Screen sharing does not guarantee it; collaborators may occupy different viewports, zoom levels, or selection states while assuming alignment.
Why it happens
Deictic language, gesture, and local edits depend on a mutually identifiable focus. Co-located partners use gaze, body orientation, pointing, and object position as reciprocal evidence. A remote interface that transmits content without focus leaves speakers unable to know whether partners have located the referent. Once focus is established, coordinates can collapse into “here”; after focus drifts, the same shorthand causes errors.
Studying it
Referential tasks can manipulate shared view, viewport coupling, pointers, or gaze cues and measure time to first location, repair turns, wrong actions, and completion time. Dialogue and screen recordings must be synchronized to determine where each partner was looking when “here” was produced. Benefits of a common view cannot all be attributed to joint attention because visual search may also be reduced.
Where it stops holding
Joint attention is not agreement or semantic understanding. Partners can inspect the same chart and reach different conclusions. Continuous gaze disclosure creates surveillance and privacy risks, and precise eye tracking is unnecessary for many tasks. Forced viewport coupling also harms autonomy during independent parallel work.
Applying it
- Expose the shared object, viewport, focus, and follow state during explanation or joint action.
- Offer a lightweight “look here” request and evidence of arrival rather than permanent gaze broadcast.
- Let followers detach and recover their prior position.
- Test with deliberately misaligned viewports and count references or actions made before the addressee locates the target.