Who reported this story?

This story was reported by arXiv cs.RO.

Robotics

Towards Generalizable Robotic Manipulation in Dynamic Environments

Robos News Newsroom

Editorial Desk

2026-07-01 · 2 min read

Published July 1, 2026 · Category: Robotics

Overview

arXiv:2603.15620v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models excel in static manipulation but struggle in dynamic environments with moving targets. This performance gap primarily stems from a scarcity of dynamic manipulation datasets and the reliance of mainstream VLAs on single-frame observations, restricting their spatiotemporal reasoning capabilities. To address this, we introduce DOMINO, a large-scale dataset and benchmark for generalizable dynamic manipulation, featuring 35 tasks with hierarchical complexities, over 110K expert trajectories, and a multi-dimensional evaluation suite. Through comprehensive experiments, we systematically evaluate existing VLAs on dynamic tasks, explore effective training strategies for dynamic awareness, and validate the generalizability of dynamic data. Furthermore, we propose PUMA, a dynamics-aware VLA architecture. By integrating scene-centric historical optical flow and specialized world queries to implicitly forecast object-centric future states, PUMA couples history-aware perception with short-horizon prediction. Results demonstrate that PUMA achieves state-of-the-art performance, yielding a 6.3% absolute improvement in success rate over baselines. Moreover, we show that training on dynamic data fosters robust spatiotemporal representations that transfer to static tasks. All code and data are available at https://github.com/H-EmbodVis/DOMINO.

Source

Originally published at arxiv.org.

Source: https://arxiv.org/abs/2603.15620

Robos News Newsroom

Robos News covers markets, crypto and commodities for Asia & the Middle East — tier-1 desk research, AI-driven analysis, institutional-grade data. Tip our newsroom: [email protected]

Email the newsroom →

Disclaimer: This article is for informational purposes only and does not constitute investment advice. Data may be delayed up to 15 minutes. Past performance is not indicative of future results. Consult a licensed financial advisor before making investment decisions.

Towards Generalizable Robotic Manipulation in Dynamic Environments

Overview

Source

Related Articles

Related Stories

Overview

Source

Related Articles

Related Stories

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation

Multi-Robot Coordination for Planning under Context Uncertainty

Hierarchical 3D Scene Graph Construction and Belief-based Planning for Semantic Navigation

LDHP: Library-Driven Hierarchical Planning for Non-prehensile Dexterous Manipulation

Cookie Preferences