← Back to blog

Redson Dev brief · PRIMARY SOURCE

ARTICLE#AI#Dev

AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR

Google Research · August 25, 2026

Integrating realistic hand gestures into Extended Reality (XR) interactions offers a significant step forward in making digital agents feel more naturally present and intuitive for users. This Google Research piece introduces AgentHands, a method for generating interactive, contextually appropriate hand gestures for virtual agents within spatially grounded XR conversations. It leverages large language models (LLMs) to infer suitable gestures from conversational content and user interactions, allowing agents to point, illustrate, or emphasize details in ways that align with human communication norms, rather than relying on pre-scripted animations or static poses. For a freelance architectural visualizer in Austin, Texas, this means moving beyond static client presentations. Instead of just showing a digital model, they could deploy an XR agent that fluidly gestures towards structural elements, points out material textures on a virtual building facade, or even indicates traffic flow patterns on an interactive city plan, making design walkthroughs far more engaging and informative for clients. Similarly, an internal IT support team at a mid-sized healthcare provider in Boston, Massachusetts, could use an XR agent with AgentHands to guide new employees through complex software interfaces. The agent might point directly at specific buttons, highlight menu options, or demonstrate multi-step processes with natural hand movements, streamlining onboarding and reducing the cognitive load on staff learning new systems. Even a small e-commerce shop based in Portland, Oregon, could deploy an XR assistant on their virtual storefront that uses natural gestures to explain product features, demonstrate scale, or recommend complementary items, elevating the shopping experience beyond simple text and images. To capitalize on this, consider how natural, contextual gestures could enhance your existing digital interfaces or services. This week, take a common interaction point in your product or service that currently relies solely on text or voice. Draft a few conversational snippets an agent might use there. Then, visualize how a virtual hand, capable of pointing, drawing shapes, or emphasizing words, could make that interaction clearer, faster, or more engaging for your users. Think about the specific spatial data or conversational context an LLM would need to infer those gestures effectively.

Source / further reading

Learn more at Google Research