AgentHands and the Limit of Voice in XR
AgentHands is not a prettier chatbot. Google’s XR prototype tests whether timed hands can finish instructions a voice-only assistant still leaves hanging.

Table of Contents
Google’s AgentHands prototype in XR does not make assistants more talkative. It tests whether a timed hand can finish the sentence a voice leaves hanging in a real room. AI assistants are getting better at talking. That crack is easy to miss on a phone. Hold up the camera, ask where the HDMI port is, and a box appears around the right socket. Look down at a plant you are supposed to water, or at a 3D printer whose knob you have never turned, and the same trick falls apart. Your hands are busy. The model can describe the motion and still leave you guessing which object, which direction, and how far.
Google’s AgentHands prototype tries to fix that mismatch in XR. It is not a shipping consumer feature. It is a CHI 2026 system that gives a conversational agent visible hands and times those hands to spoken words so an instruction can land on a real object in the room. Google posted the AgentHands write-up in August 2026; the system is also in a CHI 2026 paper. The claim worth taking seriously is narrower than “AI will have bodies.” Voice-only assistants break down the moment the task is in the room with you. AgentHands counts because it treats hands as part of the answer, not as decoration on a talking head.
Phone boxes do not survive contact with a workbench
On a phone or a pair of glasses that still think like a phone, spatial help is usually a rectangle. Project Astra-style and Gemini Live-style overlays can mark a thing in the camera frame. That is useful when the job is “find this.” It is weak when the job is “do this to that.” A bounding box can tell you which pot is the orchid. It cannot lift the pot so water drains, turn a printer knob through a click, or put a heat warning on a nozzle you should not touch while it cools.
That is where voice alone runs out of room. A paragraph like “turn the leveling dial clockwise until you feel light resistance, then back off a quarter turn” is accurate English. In the headset, the user turns the wrong dial, turns it counter-clockwise, or turns it until something snaps. The error is not hearing the sentence. The error is mapping the sentence onto physical objects while your attention is split between looking, listening, and using your hands.
Humans solved that problem by pointing and miming while they talk. You say “turn this one” and your finger points at the knob; you say “this far” and your wrist makes the motion. Language is multimodal because human speech is multimodal. When people explain physical tasks, gesture carries the geometry.
AgentHands starts from that old fact and moves it into a headset. The system builds a rough registry of objects with eye gaze and scene reconstruction, then gives the language model a way to aim a hand at those objects while it talks.
The hand is only interesting because it is tied to a word
A lot of XR demos put pretty arms on an avatar and call it embodiment. The more useful idea in AgentHands is a GestureEvent. The model does not emit a paragraph and then hope an animator guesses the motion. It writes the spoken reply with gesture instructions attached to trigger words. A local parser on the headset reads those events, lines them up with word-level timestamps from text-to-speech, and drives an animation engine so the motion lands on the syllable that needs it.
That sounds like plumbing. It is the whole product argument. If the hand arrives a second late, it becomes a distraction. If it arrives on the wrong noun, it points at the table when the sentence is about the nozzle. Co-speech timing is what keeps the gesture from turning into a second, competing UI.
Think of a cooking video where the host says “fold the whites in” while the spatula is still in the sink. You stop trusting the picture. AgentHands is trying to keep the spatula in frame when the verb happens.
The taxonomy behind those events is plain. Hands can be one or two. A palm can mean caution. A cylindrical grip can mime a tool. Motion can float in mid-air, lock to an object, or sit relative to the user. Time and visual effects do extra work: a pouring path, a tracing outline, a red glow on something hot. Interactivity shows up when the agent meets the user’s own hand, as in a warning that holds on.
Pointing, showing and warning are three different jobs
The gesture library is grouped the way people already gesture in kitchens and workshops. Deictic gestures reference. They answer “which one” and “which way.” In XR that is the difference between “the air roots” as a phrase and a hand that traces the actual roots on the plant in front of you.
Iconic gestures depict. They answer “what does the action look like.” A “turn and click” on a 3D printer is a poor sentence and a decent motion. You can hear the words and still rotate the wrong way. A hand that performs the sequence reduces the translation work your head has to do.
Expression gestures carry social and emotional cues. A wellness-coach example in the research is the least mechanical of the three, and the easiest to dismiss as theater. It still has a job: a warning that only lives in the voice is easy to file next to every other warning the model issues in a day. A warning that occupies your hand and paints a visible effect is harder to shrug off.
Those three jobs should not be mashed into one “friendly hand” animation. Pointing at the wrong object is a different failure than smiling while you describe a burn risk. If a later product blurs them, the system will look busy and still mis-teach.
Orchid pots beat keynote stages
The research demos that matter are slightly unglamorous.
In the orchid-care task, the agent goes to the base of the plant and outlines air roots while it talks. One participant said the lifting step only clicked when the agent showed the lift: they had not realized the pot needed to come up so water could drain. That sentence is the paper’s sharpest exhibit. The speech was already there. The missing piece was the motion that turned a noun into a procedure.
The 3D-printer walkthrough aims at the same class of error. Knobs, clicks, file selection, a nozzle that can burn you. A “burn” effect on the hot part is not subtle design. That is the point. A warning you can quote later is weaker than a warning you flinch from in space.
There is also a lifestyle-coach scenario, where the agent uses a warning gesture and a visual effect against a habit. Treat that as a probe, not as proof that people want an AI to grab their hand about snacks. The harder tests are still the plant and the printer, because those tasks have a correct physical outcome. You either drained the pot or you did not.
A workbench is a rude reviewer. If the hand points a few centimeters off, the user tightens the wrong fitting. Phone-box assistants can hide that error behind a wide rectangle. Spatial hands cannot.
Twelve people can show a problem and they cannot close it
The team ran a formative study with 10 people to shape the taxonomy, then a within-subjects study with 12 people comparing AgentHands with a speech-only baseline. The spoken content stayed the same. The hands did not. On a 7-point scale, participants found it easier to locate objects and directions with the hands on, and easier to follow complex actions. Both of those differences reached p < 0.05. People also described safety cues as more effective when a gesture and an effect rode along with the warning. Some talked about the hands as a partner rather than a search tool. It is also a small, lab-shaped signal.
Twelve users can tell you the speech-only condition left a mapping gap. They cannot tell you how this holds up after a week of false points, messy desks, or a model that attaches a GestureEvent to the wrong word. They also cannot tell you whether a red glow stays impressive on the fortieth warning, or what happens when two objects sit close enough that object-anchored motion becomes ambiguous.
Read the result as a diagnosis of voice, not as a verdict on Google’s XR roadmap. The speech-only baseline failed in predictable ways. Adding timed, object-aware hands reduced that failure in a short study. Shipping software lives in the long study nobody has run yet.
If Android XR only adds pretty arms, this will stall
Google frames AgentHands as a step from 2D assistant overlays toward Android XR, where the conversation is supposed to be about the room rather than a floating panel. Earlier work in that line, Human I/O and Sensible Agent, asked when a person is already busy and how an agent can avoid constant poking. Hands fit that story only if they appear when a spatial reference is required, then get out of the way.
The product fork is simple. One path glues a pair of rendered hands onto an assistant that still answers like a chatbot. The hands wave, the model talks, and users learn to ignore both. That path will photograph well and teach poorly. The other path makes gesture a referring mechanism. The model has to decide which noun needs a point, which verb needs a mime, and which warning needs a glow. It has to keep those choices aligned with a live scene that changes when the user picks up the pot. That is closer to what AgentHands is actually building: not a mascot, a timing layer between language and space.
People already know how to take instruction from hands. A colleague leans over the bench, pinches a hose clamp, and you copy the pinch. You do not ask for a paragraph. An XR assistant that can do a cheaper version of that lean-in will feel finished in a way today’s voice overlays do not. An XR assistant that talks while a pair of idle hands hover near your coffee cup will feel like a meeting with someone who will not stop gesturing at nothing. AgentHands is early, small, and still a prototype. What it gets right is the complaint users already have, even if they do not use research words for it. The assistant understood the sentence. The room did not.
Share this publication
Related Publications

Build AI Agents to Automate Film Production at Google's Agentic Cinema Hackathon
Google Cloud launched the Agentic Cinema: Blockbuster Hackathon asking developers to build Gemini-powered AI agents for film and media production with a $75,000 prize pool.

Google Is Offering Major Funding to Save Humanity’s Future — Apply Before September 25
Google opens applications for its 2026 Carbon Removal and Superpollutant Elimination R&D Awards, offering up to $500,000 per research project.