Avatar pose is the main cue to others' intent; without it, collaboration drops
Aliases: intent cues · co-presence · body orientation · head-and-hands avatar
What it is
Is the person across the table about to grab this screw, or only passing? Head and hand heading, a pause, weight shifting forward write the next intent on the body. Strip that body out of mixed reality — a name tag, a static orb — and the avatar embodiment channel is cut. Collaboration drops from seeing what the other is about to do, to guessing or asking.
Intent rides on pose, not on a realistic face. Without head and hands, realistic skin does not help.
Why it happens
Co-located work runs heavily on nonverbal lookahead: people read gaze from head heading, the next grasp from a preparatory hand, the next stance from torso orientation. Those signals beat sentences, and they interrupt the current operation less. A headset already hides the real face from bystanders in the room; if the shared scene also lacks a moving head–hands–torso, remote or semi-remote partners are left with voice. Voice queues, and it has to describe “the one on the left” — exactly the cost spatial collaboration was meant to save.
Absence also breaks turn-taking. Two people reach for the same object because neither saw the other already start. An avatar that is too latent, or simplified into a teleporting point, closes the lookahead window the same way; people fall back to “I’ll say before I move.”
Studying it
Joint assembly or sorting, comparing head+hands, head-only, name-tag-only, and no avatar, on completion time and spoken clarifications. A Networked Minds-style co-presence questionnaire can sit beside; behaviour should lead.
Conditions: avatar completeness, pose update rate, whether the real body is visible (same room vs remote). Records: task time, conflicts on the same object, “what are you looking at” questions, accuracy at predicting the partner’s next action.
When both people can see each other’s real bodies, the virtual avatar’s increment is covered. To measure the avatar itself, at least one side must not see the real body.
Where it stops holding
Purely spoken, non-spatial work (editing a document together) does not ride on pose. Two people who always face the same work surface and can see each other’s real hands find a virtual avatar redundant. Pose updates below about ten hertz, or frequent teleports, turn the avatar into interference: people explain the jump instead of reading intent. Which poses read as “I am about to step in” varies by culture; head–hand heading as gaze and reach cues still travels. Permission settings govern who may edit an object, not whether others can see that you are about to. Do not let one stand in for the other.
Applying it
- Shared manipulation should stream 6DoF of the head and both hands. Do not substitute a static badge. At least interpolate a torso heading from head position.
- On tracking loss, put the avatar in an explicit “pose unknown” (translucent, hands gone). Do not freeze the last pose — a freeze is read as “still staring at this screw.”
- Same-room with a visible real body can downgrade the avatar; a remote joiner must keep the full pose channel.
- How to check: hide the partner’s real body, leave only the shared scene, and run a two-person grab of the same part. Without head and hands, clarifications and conflicts should rise sharply; if they still ask “which one are you taking” after head and hands are on, update rate or completeness is still short.
Related
- Same group:N5.10.2 Simultaneous operation of a shared object needs an explicit control-arbitration rule · N5.10.3 Network latency briefly desynchronizes how participants see the same object's state · N5.10.4 When devices of different capability share a room, the experience must align to the weakest · N5.10.5 Accuracy of pointing in a joint task depends on precise sync of participant positions
- Nearby:N5.05 Shared Space · N5.09 Spatial Anchors and Persistence · N3.08 World-locked, Body-locked, and Head-locked
- Search terms:
avatar embodiment·nonverbal intent·co-presence
Cards in the same group
- N5.10.2Simultaneous operation of a shared object needs an explicit control-arbitration rule
- N5.10.3Network latency briefly desynchronizes how participants see the same object's state
- N5.10.4When devices of different capability share a room, the experience must align to the weakest
- N5.10.5Accuracy of pointing in a joint task depends on precise sync of participant positions