Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesLOPAL: Local Performance-Aware Active Learning from Imperfect Demonstrations
arXiv:2606.16888v1 Announce Type: new Abstract: Learning from Demonstration (LfD) enables intuitive robot skill acquisition by allowing robots to learn directly…
Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models
arXiv:2606.16902v1 Announce Type: new Abstract: This work addresses spatial question answering for service robots traversing long egocentric routes. Given a que…
EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video
arXiv:2606.16202v1 Announce Type: cross Abstract: Humans naturally understand object physics through everyday interactions, but faithfully predicting complex de…
MotionVLA: Vision-Language-Action Model for Humanoid Motion
arXiv:2606.15142v1 Announce Type: cross Abstract: Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and…
X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
arXiv:2606.14752v1 Announce Type: cross Abstract: Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise contin…
Hamilton-Jacobi Reachability-Based Safe Reinforcement Learning for Emergency Collision Avoidance
arXiv:2606.15311v1 Announce Type: cross Abstract: Emergency collision avoidance under extreme driving conditions demands safety-critical control that accounts f…
T-Rex: Tactile-Reactive Dexterous Manipulation
arXiv:2606.17055v1 Announce Type: new Abstract: The ability to react dynamically to tactile signals has long been considered crucial to agile human-level dexter…
Think Less, Act Early: Reinforced Latent Reasoning with Early Exit in Vision-Language-Action Models
arXiv:2606.15099v1 Announce Type: cross Abstract: Existing Vision-Language-Action (VLA) models predominantly rely on explicit Chain-of-Thought (CoT) reasoning t…
Beyond English: Uncovering the Multilingual Gap in Vision-Language-Action Models
arXiv:2606.15714v1 Announce Type: cross Abstract: Vision-Language-Action models have recently demonstrated promising capabilities in learning generalist robot p…
Anisotropic Template Ans\"{a}tze for Robust Positive Invariance under State-Dependent Uncertainty
arXiv:2606.16068v1 Announce Type: cross Abstract: We establish sufficient conditions for robust positive invariance under state- and input-dependent disturbance…
Distributed Safe Consensus Under Asymmetric Input and Time-Varying Output Constraints
arXiv:2606.16116v1 Announce Type: cross Abstract: This paper studies safe distributed consensus for single-integrator multi-agent systems over connected undirec…
ROSA-RL: Uncertainty-Aware Roundabout Optimized Speed Advisory with Reinforcement Learning
arXiv:2606.16558v1 Announce Type: cross Abstract: Roundabouts challenge automated driving in mixed traffic, as heterogeneous and non-deterministic human behavio…
Towards mm-Level Accurate UWB Radar: High-Accuracy Phase-Based Obstacle Detection through Multi-Channel Fusion
arXiv:2606.16657v1 Announce Type: cross Abstract: Accurate, tag-free distance estimation with ultrawideband (UWB) radar is essential for applications such as au…
ReMoBot: Retrieval-Based Few-Shot Imitation Learning for Mobile Manipulation with Vision Foundation Models
arXiv:2408.15919v4 Announce Type: replace Abstract: Imitation learning (IL) algorithms typically distill demonstrations into parametric policies to mimic expert…
Explainable deep learning improves human mental models of self-driving cars
arXiv:2411.18714v3 Announce Type: replace Abstract: Self-driving cars increasingly rely on deep neural networks to achieve human-like driving. The opacity of su…
AC-LIO: Towards Asymptotic Compensation for Distortion in LiDAR-Inertial Odometry via Selective Intra-Frame Smoothing
arXiv:2412.05873v4 Announce Type: replace Abstract: Existing LiDAR-Inertial Odometry (LIO) methods typically utilize the prior trajectory derived from the IMU i…
Intelligent Sailing Model for Open Sea Navigation
arXiv:2501.04988v2 Announce Type: replace Abstract: Autonomous vessels potentially enhance safety and reliability of seaborne trade. To facilitate the developme…
DemoDiffusion: One-Shot Human Imitation using pre-trained Diffusion Policy
arXiv:2506.20668v3 Announce Type: replace Abstract: We propose DemoDiffusion, a simple method for enabling robots to perform manipulation tasks by imitating a s…
Latent Action Pretraining Through World Modeling
arXiv:2509.18428v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have gained popularity for learning robotic manipulation tasks that foll…
C-3TO: Continuous 3D Trajectory Optimization on Neural Euclidean Signed Distance Fields
arXiv:2509.20084v2 Announce Type: replace Abstract: This paper introduces a novel framework for continuous 3D trajectory optimization in cluttered environments,…
RSPECT: Robust and Scalable Planner for Energy-Aware Coordination of UAV-UGV Teams in Aerial Monitoring
arXiv:2511.21957v2 Announce Type: replace Abstract: We consider the robust planning of energy-constrained unmanned aerial vehicles (UAVs) and unmanned ground ve…
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
arXiv:2601.04061v2 Announce Type: replace Abstract: Generalist Vision-Language-Action models remain constrained by the scarcity of robotic data relative to the …
LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
arXiv:2601.05248v4 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking …
Neural Minimum-Distance Estimation for Collision-Aware Operation of Multi-Arm Laparoscopy Surgical Robots Through Learning-from-Simulation
arXiv:2601.15459v2 Announce Type: replace Abstract: This study presents an integrated framework for enhancing the safety and operational efficiency of robotic a…
IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance
arXiv:2601.16207v2 Announce Type: replace Abstract: Many Vision-Language-Action (VLA) models flatten image patches into a 1D token sequence, weakening the 2D sp…
Bimanual High-Density EMG Control for In-Home Mobile Manipulation by Users with Quadriplegia
arXiv:2602.02773v2 Announce Type: replace Abstract: Mobile manipulators in the home can enable people with cervical spinal cord injury (cSCI) to perform daily p…
HiCrowd: Hierarchical Crowd Flow Alignment for Dense Human Environments
arXiv:2602.05608v3 Announce Type: replace Abstract: Navigating through dense human crowds remains a significant challenge for mobile robots. A key issue is the …
Activity-Dependent Plasticity in Morphogenetically-Grown Recurrent Networks
arXiv:2604.03386v2 Announce Type: replace Abstract: Developmental approaches to neural architecture search grow functional networks from compact genomes through…
Human Cognition in Machines: A Unified Perspective of World Models
arXiv:2604.16592v2 Announce Type: replace Abstract: This report of world models distinguishes prior works by the cognitive functions they innovate. Many works c…
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
arXiv:2605.27284v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also fol…