Who reported this story?

This story was reported by arXiv cs.RO.

Robotics

LARA: Latent Action Representation Alignment for Vision-Language-Action Models

Robos News Newsroom

Editorial Desk

2026-07-01 · 2 min read

Published July 1, 2026 · Category: Robotics

Overview

arXiv:2606.07100v2 Announce Type: replace-cross Abstract: Visual-language action (VLA) models enable robots to predict actions directly from observations and language instructions, but their performance depends on large-scale, high-quality data and is limited by the scarcity of real-world robot action datasets. To facilitate VLA model learning with abundant unlabeled human videos, Latent Action Models (LAM) learn latent action representations from visual dynamics to provide additional supervision for VLA learning. However, LAM and VLA are typically trained separately, leaving LAM ungrounded during VLA training and VLA models constrained by frozen LAM representations. To address these issues, we propose Latent Action Representation Alignment (LARA), a plug-and-play framework that jointly optimizes LAM and VLA via representation alignment. This enables reciprocal benefits where LAMs learn with action trajectories to avoid spurious visual changes, while VLAs are regularized by forward dynamics learned within LAMs to reduce hallucinations of functionally ineffective trajectories. We demonstrate LARA versatility and effectiveness for pre-training, post-training enhancement of pre-trained VLA models, and LAM refinement, achieving an average of ~10%, ~5%, and ~15% improvement over 3 simulation and 1 meticulously designed real-world robotic manipulation benchmarks.

Source

Originally published at arxiv.org.

Source: https://arxiv.org/abs/2606.07100

Robos News Newsroom

Robos News covers markets, crypto and commodities for Asia & the Middle East — tier-1 desk research, AI-driven analysis, institutional-grade data. Tip our newsroom: [email protected]

Email the newsroom →

Disclaimer: This article is for informational purposes only and does not constitute investment advice. Data may be delayed up to 15 minutes. Past performance is not indicative of future results. Consult a licensed financial advisor before making investment decisions.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models

Overview

Source

Related Articles

Related Stories

Overview

Source

Related Articles

Related Stories

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation

Multi-Robot Coordination for Planning under Context Uncertainty

Hierarchical 3D Scene Graph Construction and Belief-based Planning for Semantic Navigation

LDHP: Library-Driven Hierarchical Planning for Non-prehensile Dexterous Manipulation

Cookie Preferences