Field Notes
The Assistant Learned to Point
A virtual hand can make spatial guidance easier to follow. It can also make an uncertain answer feel like a fact already located in the room.
The assistant tells you to watch the hot part of the 3D printer. A blue hand appears beside the machine, points at the nozzle, and gives a small warning gesture as the explanation reaches the word hot.
The sentence no longer has to carry the whole instruction. You do not have to translate “the assembly above the print bed” into a hurried search across an unfamiliar machine. The answer has entered the room.
Google Research recently described AgentHands, an experimental system that gives a conversational assistant animated hands in extended reality. The hands can point to a registered object, trace a shape, demonstrate a movement, mark a distance, signal caution, or offer a celebratory high-five. They move in time with the assistant’s speech.
This sounds like ornament until you try to explain a physical task without being able to point.
Anyone who has talked a relative through a cable problem over the phone knows the strange labor of converting a room into sentences. The port on the left. Whose left? The small black cable. There are four. Turn it gently. Around which axis, and how gently is gently? Language keeps reaching for a shared scene it cannot quite touch.
AgentHands closes part of that gap with a specific mechanism. A person looks at an object and asks the system to register it. Eye gaze intersects with a reconstructed scene; an image crop and spatial coordinates go to a multimodal model; the object returns with a label and a three-dimensional bounding box. When the assistant later speaks, its response includes timed gesture events attached to particular words. An animation system turns those events into a hand that points, pours, measures, or pauses in the relevant place.
In the CHI paper, twelve participants followed scripted orchid-care and 3D-printer instructions with and without the virtual hands. They found locations, actions, and warnings easier to understand with the gestures. The study is small, the spoken content was held constant, and the environments were deliberately prepared. It establishes that the hand helped people follow these particular demonstrations. It does not establish that the assistant’s underlying advice was right.
That distinction becomes more important when an answer can point.
Text leaves a little distance between a claim and the world. A sentence saying “use this valve” still asks the reader to perform the mapping. A hand hovering beside one valve completes the mapping on the system’s behalf. The gesture is useful because it reduces ambiguity. It is persuasive for the same reason.
The finger is another claim.
It claims that the object registry is current, that gaze selected what the person intended, that the bounding box surrounds the correct part, that the gesture arrived on the correct word, and that the instruction deserves to become action. None of those claims is visible in the elegance of the movement. A perfectly synchronized hand can point beautifully at the wrong screw.
The researchers are candid about some of this. Their current object registry depends on a lightweight pre-scan and works best in relatively static environments. Fine-grained work such as circuit assembly may require target snapping or more precise cues. The gesture library is manually curated. Future versions may remember a person’s room and routines, which could make the system more helpful while giving a stale spatial assumption a much longer life.
This is where embodied assistance needs the equivalent of draft marks.
The interface should show what object the system believes it has registered before the hand begins teaching. If the scene changes, the pointer should lose confidence visibly rather than glide toward an old location. A person should be able to ask “why that part?” and receive the image, label, or spatial relation that grounded the gesture. Consequential guidance should keep the spoken instruction close enough to inspect after the hand has disappeared.
The alternative path matters too. The W3C’s XR accessibility requirements emphasize accurate object descriptions, multiple input methods, motion-agnostic interaction, adjustable cues, and ways to mute non-critical animation. A gesture can reduce cognitive load for one person and become inaccessible clutter for another. The hand should be one expression of the guidance, not the only place where meaning lives.
This extends the argument made in The Same Question, A Different Assistant: an assistant’s role can change without its factual content changing. There, language altered how readily the system sounded warm, rigorous, candid, or deferential. Here, nonverbal behavior changes how present and actionable the answer feels. A cautionary sentence accompanied by an open palm carries different social force. A high-five can turn task completion into approval. A point can make uncertainty feel settled.
That force is not a reason to keep assistants trapped in chat boxes. Spatial guidance can be genuinely humane. It can spare someone from holding ten verbal steps in memory while both hands are occupied. It can make an unfamiliar tool less forbidding. It can let an explanation meet a person where the confusion actually is: in front of the machine, not inside a paragraph.
But once the assistant learns to point, interface quality includes the honesty of the gesture. The system is no longer only choosing words. It is placing confidence in the room.
The useful hand will not merely show us where to look. It will help us see what it thinks it is pointing at, how certain that mapping is, and how to continue when the hand cannot be trusted.