Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesIVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance
arXiv:2601.16207v2 Announce Type: replace Abstract: Many Vision-Language-Action (VLA) models flatten image patches into a 1D token sequence, weakening the 2D sp…
Bimanual High-Density EMG Control for In-Home Mobile Manipulation by Users with Quadriplegia
arXiv:2602.02773v2 Announce Type: replace Abstract: Mobile manipulators in the home can enable people with cervical spinal cord injury (cSCI) to perform daily p…
HiCrowd: Hierarchical Crowd Flow Alignment for Dense Human Environments
arXiv:2602.05608v3 Announce Type: replace Abstract: Navigating through dense human crowds remains a significant challenge for mobile robots. A key issue is the …
Activity-Dependent Plasticity in Morphogenetically-Grown Recurrent Networks
arXiv:2604.03386v2 Announce Type: replace Abstract: Developmental approaches to neural architecture search grow functional networks from compact genomes through…
Human Cognition in Machines: A Unified Perspective of World Models
arXiv:2604.16592v2 Announce Type: replace Abstract: This report of world models distinguishes prior works by the cognitive functions they innovate. Many works c…
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
arXiv:2605.27284v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also fol…
WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation
arXiv:2606.04907v2 Announce Type: replace Abstract: Visual navigation requires generating smooth and collision-free trajectories under complex geometric and phy…
Efficient Reinforcement Learning by Guiding World Models with Non-Curated Data
arXiv:2502.19544v3 Announce Type: replace-cross Abstract: Leveraging offline data is a promising way to improve the sample efficiency of online reinforcement le…
CropTrack: A Tracking with Re-Identification Framework for Precision Agriculture
arXiv:2512.24838v2 Announce Type: replace-cross Abstract: Multiple-object tracking (MOT) in agricultural environments presents major challenges due to repetitiv…
The embodied brain: Bridging the brain, body, and behavior with biorealistic neuromechanical models
arXiv:2601.08056v3 Announce Type: replace-cross Abstract: Animal behavior reflects interactions between the nervous system, body, and environment. Therefore, bi…
Learning Fine-Grained Correspondence with Cross-Perspective Perception for Open-Vocabulary 6D Object Pose Estimation
arXiv:2601.13565v2 Announce Type: replace-cross Abstract: Open-vocabulary 6D object pose estimation empowers robots to manipulate arbitrary unseen objects guide…
Seeing Roads Through Words: A Language-Guided Framework for RGB-T Driving Scene Segmentation
arXiv:2602.07343v2 Announce Type: replace-cross Abstract: Robust semantic segmentation of road scenes under adverse illumination, lighting, and shadow condition…
ROSA: Roundabout Optimized Speed Advisory with Multi-Agent Trajectory Prediction in Multimodal Traffic
arXiv:2602.14780v2 Announce Type: replace-cross Abstract: We present ROSA -- Roundabout Optimized Speed Advisory -- a system that combines multi-agent trajector…
Systematic Evaluation of Novel View Synthesis for Video Place Recognition
arXiv:2603.05876v2 Announce Type: replace-cross Abstract: The generation of synthetic novel views has the potential to positively impact robot navigation in sev…
Safe Exploration via Policy Priors
arXiv:2601.19612v3 Announce Type: replace-cross Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online…
Artists' Views on Robotics Involvement in Painting Productions
arXiv:2510.07063v3 Announce Type: replace-cross Abstract: As robotic technologies evolve, their potential in artistic creation becomes an increasingly relevant …
Bio-inspired decision making in robot swarms under biases
arXiv:2509.07561v2 Announce Type: replace-cross Abstract: Minimalistic robot swarms offer a scalable, robust, and cost-effective approach to performing complex …
DynNPC: Finding More Violations Induced by ADS in Simulation Testing through Dynamic NPC Behavior Generation
arXiv:2411.19567v3 Announce Type: replace-cross Abstract: Recently, a number of simulation testing approaches have been proposed to generate diverse driving sce…
Imitating What Works: Simulation-Filtered Modular Policy Learning from Human Videos
arXiv:2602.13197v2 Announce Type: replace Abstract: The ability to learn manipulation skills by watching videos of humans has the potential to unlock a new sour…
Simplifying ROS2 controllers with a modular architecture for robot-agnostic reference generation
arXiv:2601.08514v2 Announce Type: replace Abstract: This paper introduces a novel modular architecture for ROS2 that decouples the logic required to acquire, va…
Multi-Robot Motion Planning from Vision and Language using Heat-Inspired Diffusion
arXiv:2512.13090v2 Announce Type: replace Abstract: Diffusion models have recently emerged as powerful tools for robot motion planning by capturing the multi-mo…
Bayesian Optimization for Learning Nonlinear MPC in Autonomous Agent Navigation
arXiv:2606.14763v1 Announce Type: new Abstract: Real-time autonomous navigation in dynamic, unknown environments remains a fundamental challenge for mobile robo…
OmniVTLA: Vision-Tactile-Language-Action Models with Semantic-Aligned Tactile Sensing
arXiv:2508.08706v3 Announce Type: replace Abstract: Recent vision-language-action (VLA) models build upon vision-language foundations, and have achieved promisi…
Direction-Conditioned Policies via Compositional Subgoal Scoring for Online Goal-Conditioned Reinforcement Learning
arXiv:2606.16515v1 Announce Type: cross Abstract: Hamilton-Jacobi-Bellman theory implies that the optimal goal-conditioned action depends on the goal only throu…
MVOFormer: Flow-Semantic Transformer for Robust Monocular Visual Odometry
arXiv:2606.16474v1 Announce Type: cross Abstract: Monocular visual odometry (MVO) is foundational to autonomous navigation and robotic localization. However, ex…
FlowMPC: Improving Flow Matching policies with World Models
arXiv:2606.16286v1 Announce Type: cross Abstract: Flow Matching (FM) is a powerful approach for behavior cloning in multimodal action spaces [Jiang et al., 2025…
Towards Next-Generation Healthcare: A Survey of Medical Embodied AI for Perception, Decision-Making, and Action
arXiv:2606.15647v1 Announce Type: cross Abstract: Foundation models have demonstrated impressive performance in enhancing healthcare efficiency across a wide ra…
Robust Conformal CBF and CLF Controllers via Iterative Policy Updates
arXiv:2606.15366v1 Announce Type: cross Abstract: Conformal prediction (CP) has been used to obtain probabilistic bounds on the error between a learned dynamics…
VLALeaks: Membership Inference Attacks against Vision-Language-Action Models
arXiv:2606.15165v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models enable end-to-end robot control and have garnered widespread attention. Ho…
Phase-Localized Curation Does Not Help: A Negative Result on Per-Phase Metric Selection for Demonstration Filtering
arXiv:2606.15064v1 Announce Type: cross Abstract: Manipulation demonstrations have temporal phase structure, and a natural hypothesis is that demonstration-cura…