Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 stories3D Scene Graphs: Open Challenges and Future Directions
arXiv:2606.19383v1 Announce Type: new Abstract: 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric groundin…
Temporal Self-Imitation Learning
arXiv:2606.19752v1 Announce Type: new Abstract: Long-horizon robot manipulation policies trained with reward shaping can still exploit dense rewards through ine…
Physical Atari: A Robust and Accessible Platform for Real-time Reinforcement Learning on Robots
arXiv:2606.19357v1 Announce Type: new Abstract: We built a robot called the Robotroller that actuates an Atari CX40+ controller and a device called the Atari De…
Efficiently Linking Real Scenes with Synthetic Data Generation for AI-based Cognitive Robotics and Computer Vision Applications
arXiv:2606.20272v1 Announce Type: new Abstract: AI vision models are a driving factor for the potential use case scenarios of cognitive robotics within in the i…
Dual-Agent Framework for Cross-Model Verified Translation of Natural-Language Protocols into Robotic Laboratory Platform
arXiv:2606.20120v1 Announce Type: new Abstract: Biological experiment protocols are written in natural language, whereas automation systems rely on predefined c…
Frequency-Aware Flow Matching for Continuous and Consistent Robotic Action Generation
arXiv:2606.20135v1 Announce Type: new Abstract: Flow matching has emerged as a standard paradigm for robotic manipulation owing to its strong expressive power f…
WorkBenchMark: A LEGO-Based Assembly Benchmark with an Assembly-by-Disassembly Baseline for the Smart Manufacturing League
arXiv:2606.19358v1 Announce Type: new Abstract: We introduceWorkBenchMark, a LEGO Duplo-based robotic assembly benchmark motivated by the RoboCup Smart Manufact…
DiffusionVS: A Generative Framework for Robust Visual Servoing Based on Diffusion Policy
arXiv:2606.19397v1 Announce Type: new Abstract: Visual servoing is a fundamental technique in robotic manipulation and navigation. Regression-based visual servo…
Simulating Robotic Locomotion in Sand: Resistive Force Theory in an Open-Source Physics Engine
arXiv:2606.19504v1 Announce Type: new Abstract: Recent advancements in Resistive Force Theory (RFT) enable approximation of ground reaction forces for locomotio…
Formal Verification of Learned Multi-Agent Communication Policies via Decision Tree Distillation
arXiv:2606.19632v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent com…
Scaling Self-Play for End-to-End Driving
arXiv:2606.19641v1 Announce Type: new Abstract: End-to-end autonomous driving models are typically trained on offline human-demonstration datasets that provide …
A Differentiable Composite Approximation Framework for Autonomous Underwater Vehicle Maneuvering Modeling from Sea-Trial Data
arXiv:2606.19711v1 Announce Type: new Abstract: Field-based modeling from onboard measurements can produce autonomous underwater vehicle (AUV) maneuvering model…
Safe Local Navigation for Ackermann-Steered Robots in Unmapped Environments
arXiv:2606.19672v1 Announce Type: new Abstract: A control framework is proposed for safe local navigation of mobile robots equipped with Ackermann steering in u…
VOiLA: Vectorized Online Planning with Learned Diffusion Model for POMDP Agents
arXiv:2606.19729v1 Announce Type: new Abstract: Planning under uncertainty is an essential capability for autonomous robots. The Partially Observable Markov Dec…
Data Standards for Humanoid Robotics: The Missing Infrastructure for Physical AI
arXiv:2606.19769v1 Announce Type: new Abstract: The scalability of humanoid robots will depend not only on models and hardware, but also on whether physical exp…
Start Right, Arrive Right: Asynchronous Execution via Initial Noise Selection
arXiv:2606.19774v1 Announce Type: new Abstract: Action chunking enables robot policies to produce temporally coherent behavior, but generating multi-step action…
One-to-Two Acting: A Novel Framework for Single-arm Agent Action Expansion to Dual Arms
arXiv:2606.19897v1 Announce Type: new Abstract: Dual-arm manipulation can improve throughput via parallel execution, but collecting bimanual demonstrations for …
Co-policy: Responsive Human-Robot Co-Creation for Musical Performances
arXiv:2606.19914v1 Announce Type: new Abstract: Art has long stood as a pivotal expression of human creativity. Embodied artificial intelligence offers a route …
Fast Human Attention Prediction for Fixation-guided Active Perception in Autonomous Navigation
arXiv:2606.20491v1 Announce Type: new Abstract: Human visual attention relies on structured scanpaths to efficiently process scenes, yet instilling this behavio…
MemoryWAM: Efficient World Action Modeling with Persistent Memory
arXiv:2606.20562v1 Announce Type: new Abstract: Robust robotic manipulation in the real world requires not only an understanding of the current observation, but…
Deep-Unfolded Coordination
arXiv:2606.19920v1 Announce Type: new Abstract: Distributed optimization is a highly scalable and structurally transparent technique to solve multi-agent roboti…
Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory
arXiv:2606.19998v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are increasingly deployed across diverse tasks, yet they remain black boxes …
GroundControl: Anticipating Navigation Failures in Vision-Language Agents via Trajectory-Consistent Uncertainty Estimates
arXiv:2606.20479v1 Announce Type: new Abstract: Vision-language navigation agents achieve competitive average success on benchmark tasks, yet failures often ari…
A Neuromorphic Reinforcement Learning Framework for Efficient Pathfinding in Robotic Mobile Fulfillment Systems
arXiv:2606.20031v1 Announce Type: new Abstract: Dynamic environmental changes, confined workspaces, and stringent real-time constraints make pathfinding in Robo…
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
arXiv:2606.19531v1 Announce Type: cross Abstract: World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control…
Slow Brain, Fast Planner: Latency-Resilient VLM-Augmented Urban Navigation
arXiv:2606.20458v1 Announce Type: new Abstract: Learning-based planners for sidewalk navigation can generate diverse candidate trajectories in real time, yet th…
MirrorDuo: Reflection-Consistent Visuomotor Learning from Mirrored Demonstration Pairs
arXiv:2606.20048v1 Announce Type: new Abstract: Image-based behaviour cloning leverages demonstrations captured from ubiquitous RGB cameras. However, it remains…
Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online
arXiv:2509.00271v2 Announce Type: replace Abstract: We introduce a novel History-Aware VErifier (HAVE) to disambiguate uncertain scenarios online by leveraging …
ARC: Adaptive Robust Joint State and Covariance Estimation
arXiv:2606.20428v1 Announce Type: new Abstract: Sensor measurements are frequently corrupted by outliers and non-Gaussian noise. These imperfections in the sens…
ForEnt: A Multi-Modal Dataset for Characterizing Quadruped Robot Entrapments in Forest Environments
arXiv:2606.19675v1 Announce Type: new Abstract: Legged robots are increasingly deployed in forests for ecological surveying and monitoring, yet their autonomy i…