Voice and head-worn displays are primary alternatives
Aliases: head-worn display · voice-driven hands-free interaction · near-eye display
What it is
Voice and head-worn interaction uses speech to carry commands or spoken notes, and a head-worn display to bring information into the operator's line of sight, cutting down the hand traffic between tool and terminal. The two are often bundled together as "the hands-free solution," but they solve different kinds of problems: voice is an input channel that replaces pressing or touching; a head-worn display is mainly an output and information-placement channel that replaces looking down at a screen or flipping through paper. They can be combined, but neither substitutes for the other's function — fixing the voice problem does not deliver the display's positioning benefit, and vice versa.
Why it happens
Voice's value is that it lets an operator summon information or issue a command without releasing a grip; a head-worn display's value is shortening the visual path between equipment, drawings, and terminal, cutting the head-down/head-up switching and neck-posture load. But both consume resources that would otherwise serve the primary task: voice occupies auditory attention and the verbal channel, competing with the task whenever spoken coordination or environmental listening matters; a head-worn display occupies part of the visual field, and its frame, the refocusing cost between the near-eye display and the far scene, and its weight and balance all interfere with observing the real environment. If overlaid information sits fixed at the center of the field of view for too long, it can also produce attention tunneling — the display keeps pulling attention while peripheral real-world hazard cues go unnoticed. Speech recognition itself depends on an acoustic model adapting to the user's speech rate, accent, and ambient noise; a misrecognition carries none of the physical boundary cues a missed touch does, so it needs its own feedback mechanism before the operator can trust that a command landed correctly.
Studying it
Compare a handheld terminal, voice-only, monocular, and binocular head-worn options under the target task and environment, measuring information-lookup time, duration and frequency of gaze leaving the real scene, false-command rate, correction cost, postural load, and the quality of the primary task itself, not just the speed of the assistive task. Head-worn devices show a pronounced novelty effect — first-use excitement inflates subjective satisfaction — so evaluation should extend across repeated use and multiple shifts, with participants wearing the target PPE, coordinating with others, and moving along real paths; short, solo, stationary tests rarely predict sustained field performance.
Where it stops holding
Work that depends on peripheral vigilance for hazards, fine depth judgment (aligning precision parts), or long continuous wear may not suit a head-worn display — the cost of occluded field of view and refocusing can outweigh the benefit of positioned information. Voice does not suit high noise, a respirator mask that distorts speech, or content sensitive enough to require confidentiality. Hazardous commands should never execute purely on an open-vocabulary recognition result, because speech misrecognition has none of the physical boundary cues that a missed touch does.
Applying it
- Keep head-worn content short and directly tied to the current action, and let the operator dismiss it or switch it to transparent/hidden at will — never let it occlude the field of view over a hazard source.
- For voice commands, echo back the recognized object, action, and consequence, require a separate confirmation action rather than another spoken repetition, and support quick correction without re-running the whole command sequence.
- Provide a fallback for voice and for the head-worn display that does not depend on the other (voice-only without the display, or display-only with manual paging without voice), and validate comfort and impact on the primary task over full shifts in target PPE, not just single-task completion time.