Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesAsyncMDE: Real-Time Monocular Depth Estimation via Asynchronous Spatial Memory
arXiv:2603.10438v2 Announce Type: replace Abstract: Foundation-model-based monocular depth estimation offers a viable alternative to active sensors for robot pe…
X-Morph: Human Motion Priors for Scalable Robot Learning Across Morphologies
arXiv:2606.30290v1 Announce Type: new Abstract: Recent progress in humanoid behavior models has been driven in large part by abundant human motion data, but com…
The Speedup Paradox: Rethinking Inference Speed-Quality Trade-off in Embodied Tasks
arXiv:2606.28529v1 Announce Type: new Abstract: Embodied foundation models have recently been widely used to improve robot generalization and task success rates…
You Only Touch Once: 6-DoF Object Pose Estimation from Single Tactile Contact
arXiv:2606.28899v1 Announce Type: new Abstract: Accurate 6-DoF object pose estimation is fundamental to robotic manipulation, yet vision-based methods often fai…
Learning to Throw: Agile and Accurate Cable-Suspended Payload Delivery with a Quadrotor
arXiv:2606.27603v1 Announce Type: new Abstract: Quadrotors offer the agility needed to rapidly transport suspended payloads during time-critical applications, i…
PPO-EAL: Exact Augmented Lagrangian Proximal Policy Optimization for Safe Robotic Control
arXiv:2606.27861v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a promising solution to accomplish complex robotic control tasks; how…
On dynamic multi-agent pathfinding methods: review, simulations and modifications
arXiv:2606.03735v2 Announce Type: replace-cross Abstract: This paper presents a systematic study of pathfinding algorithms in the context of Dynamic Multi-Agent…
AI-Driven Synthesis for High-Tech System Design: Automating Innovation
arXiv:2606.28126v1 Announce Type: cross Abstract: This article addresses the combinatorial complexity inherent in modern high-tech system design by presenting a…
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
arXiv:2510.01642v3 Announce Type: replace Abstract: Recent advances in robotic manipulation have integrated low-level robotic control into Vision-Language Model…
Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots
arXiv:2606.28133v1 Announce Type: new Abstract: We study whether we can learn novel manipulation skills from human actions to a bi-manual robot with parallel gr…
Drifting in the Future: Stabilizing Path Following Drifting on High-Latency Vehicle Systems
arXiv:2606.27914v1 Announce Type: new Abstract: Autonomously controlling and handling a vehicle at and beyond its stability limit is a mathematically and comput…
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
arXiv:2512.21970v2 Announce Type: replace Abstract: While Vision-Language-Action (VLA) models excel in generalist manipulation, they often lack fine-grained spa…
Radar Guided Camera Verification for Automatic Emergency Braking Rethinking Object Detection in Radar Camera Fusion
arXiv:2606.27556v1 Announce Type: cross Abstract: Radar camera fusion is widely used in Automatic Emergency Braking AEB systems because radar provides reliable …
CWI: Composite Humanoid Whole-Body Imitation System for Loco-manipulation
arXiv:2606.27676v1 Announce Type: new Abstract: Achieving everyday tasks with humanoid robots requires coordinating stable locomotion with versatile manipulatio…
KISS-IMU: Self-supervised Inertial Odometry with Motion-balanced Learning and Uncertainty-aware Inference
arXiv:2603.06205v2 Announce Type: replace Abstract: Inertial measurement units (IMUs), which provide high-frequency linear acceleration and angular velocity mea…
Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors
arXiv:2606.28237v1 Announce Type: new Abstract: Quadruped robots have achieved remarkable locomotion, yet their behavioral repertoire remains confined to a few …
Data Scaling Laws in Imitation Learning for Robotic Manipulation
arXiv:2410.18647v4 Announce Type: replace Abstract: Data scaling has revolutionized fields like natural language processing and computer vision, providing model…
MobileManiBench: Simplifying Model Verification for Mobile Manipulation
arXiv:2602.05233v2 Announce Type: replace Abstract: Vision-language-action models have advanced robotic manipulation but remain constrained by reliance on the l…
Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?
arXiv:2606.27755v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models enable instruction-driven robotic manipulation, but they inherit oversized l…
S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation
arXiv:2606.27872v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation, but their per…
AeroGrab: A Unified Framework for Aerial Grasping in Cluttered Environments
arXiv:2603.15097v2 Announce Type: replace Abstract: Reliable aerial grasping in cluttered environments remains challenging due to occlusions and collision risks…
DIM-WAM: World-Action Modeling with Diverse Historical Event Memory
arXiv:2606.27677v1 Announce Type: new Abstract: World-action models have shown promising robot-manipulation performance by jointly predicting future visual stat…
Relating Reinforcement Learning to Dynamic Programming-Based Planning
arXiv:2603.07844v2 Announce Type: replace Abstract: This paper bridges some of the gap between optimal planning and reinforcement learning (RL), both of which s…
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
arXiv:2606.28128v1 Announce Type: cross Abstract: Video generation models have emerged as a promising paradigm for embodied world simulation. However, both gene…
Spacecraft Fiducial Marker for Autonomous Rendezvous, Proximity Operations, and Docking
arXiv:2606.27566v1 Announce Type: new Abstract: Robotic operations in space are challenging due to the harsh environment and the high cost of failure. Fiducial …
VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization
arXiv:2507.17455v2 Announce Type: replace-cross Abstract: Geo-localization from a single image at planet scale (essentially an advanced or extreme version of th…
RAE-NWM: Navigation World Model in Dense Visual Representation Space
arXiv:2603.09241v2 Announce Type: replace-cross Abstract: Visual navigation requires agents to reach goals in complex environments through perception and planni…
DIVER: Reinforced Diffusion Breaks Imitation Bottlenecks in End-to-End Autonomous Driving
arXiv:2507.04049v5 Announce Type: replace-cross Abstract: Most end-to-end autonomous driving methods rely on imitation learning from single expert demonstration…
Image-based Geo-localization for Robotics: Are Black-box Vision-Language Models there yet?
arXiv:2501.16947v2 Announce Type: replace-cross Abstract: The advances in Vision-Language models (VLMs) offer exciting opportunities for robotic applications in…
Learn Structure, Adapt on the Fly: Multi-Scale Residual Learning and Online Adaptation for Aerial Manipulators
arXiv:2603.11638v2 Announce Type: replace Abstract: Autonomous Aerial Manipulators (AAMs) are inherently coupled, nonlinear systems that exhibit nonstationary a…