Robotics

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos

Robos News Newsroom

Editorial Desk

2026-06-18 · 2 min read

Published June 18, 2026 · Category: Robotics

Overview

arXiv:2606.18955v1 Announce Type: cross Abstract: Training generalist Vision-Language-Action(VLA) models typically requires massive, diverse robotic datasets with high-fidelity action annotations. While egocentric human manipulation videos are abundant and capture significant environmental diversity, the absence of action labels makes them difficult to use in conventional training paradigms. To address this, we propose a latent-action-based framework designed to extract general action priors from unlabeled human videos. The architecture features a Hybrid Disentangled VQ-VAE that decouples motion dynamics from environmental backgrounds through physical masks, enabling the construction of a cross-embodiment action codebook. By pre-training on human videos with the codebook, the VLM backbone learns deep representations of action intent. For adaptation to specific embodiments, we introduce an intent-perception decoupling strategy where the VLM predicts the action intent while a separate frozen visual encoder provides state-specific features to the action expert, thereby reducing action hallucinations. Results in simulation and real-world environments show that our method, pre-trained exclusively on unlabeled human videos, performs competitively with state-of-the-art VLA models trained on massive annotated datasets, requiring only 50 trajectories for downstream adaptation.

Source

Originally published at arxiv.org.

Source: https://arxiv.org/abs/2606.18955

Robos News Newsroom

Robos News reports on robotics research, components, manufacturers, field deployments, and industrial automation worldwide. Tip our newsroom: [email protected]

Email the newsroom →

Reporting standard: Product specifications, deployment counts, and performance claims are attributed to their source. Safety-critical decisions should be based on the applicable technical documentation and validation for the operating environment.

Cookie Preferences

Overview

Source

Related Articles

Related Stories

Top 10 robotics stories of July 2026

KUKA deploys Automation Management Platform for North American automakers

FCC robot ruling shines a spotlight on U.S. policy; how next-gen AI can help warehousing

Procore Technologies acquires DroneDeploy for $845M

Cookie Preferences