Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesDIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?
arXiv:2606.12402v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly deployed as high-level planners for embodied agents, with an emer…
World Pilot: Steering Vision-Language-Action Models with World-Action Priors
arXiv:2606.12403v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models inherit semantic grounding from large-scale pretraining and perform competen…
VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving
arXiv:2606.12396v1 Announce Type: cross Abstract: Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle …
Non-Equilibrium MAV-Capture-MAV via Time-Optimal Planning and Reinforcement Learning
arXiv:2503.06578v2 Announce Type: replace Abstract: The capture of flying MAVs (micro aerial vehicles) has garnered increasing research attention due to its int…
Fourier Features Let Agents Learn High Precision Policies with Imitation Learning
arXiv:2606.12334v1 Announce Type: cross Abstract: High-precision robotic manipulation requires fine-grained spatial reasoning that is often difficult to achieve…
iPack: Intuitive Bin Packing with Large Language Models
arXiv:2503.08445v2 Announce Type: replace Abstract: Robotics and automation are increasingly influential in logistics but remain largely confined to traditional…
SR-LIO++: LiDAR-Inertial Odometry and Quantized Mapping with Caching-Aware Sweep Reconstruction
arXiv:2503.22926v3 Announce Type: replace Abstract: Addressing the inherent low acquisition frequency limitation of 3D LiDAR to achieve high-frequency output ha…
LEMON-Mapping: Loop-Enhanced Large-Scale Multi-Session Point Cloud Merging and Optimization for Globally Consistent Mapping
arXiv:2505.10018v4 Announce Type: replace Abstract: Multi-robot collaboration is becoming increasingly critical and presents significant challenges in modern ro…
CU-Multi: A Dataset for Multi-Robot Collaborative Perception
arXiv:2509.19463v2 Announce Type: replace Abstract: A central challenge for multi-robot systems is fusing independently gathered perception data into a unified …
Learning Ordinal Response Policies in Rank-Based Stochastic Prize-Collecting Games
arXiv:2510.24515v2 Announce Type: replace Abstract: The Team Orienteering Problem (TOP) generalizes many real-world multi-agent scheduling and routing tasks tha…
SIL: Symbiotic Interactive Learning for Language-Conditioned Human-Agent Co-Adaptation
arXiv:2511.05203v3 Announce Type: replace Abstract: Today's autonomous agents, largely driven by foundation models (FMs), can understand natural language instru…
DynaRetarget: Dynamically-Feasible Retargeting using Sampling-Based Trajectory Optimization
arXiv:2602.06827v3 Announce Type: replace Abstract: In this paper, we introduce DynaRetarget, a complete pipeline for retargeting human motions to humanoid cont…
Consensus-based optimization (CBO): Towards Global Optimality in Robotics
arXiv:2602.06868v2 Announce Type: replace Abstract: Zero-order optimization has recently received significant attention for designing optimal trajectories and p…
Vision-Aided Relative State Estimation for Approach and Landing on a Moving Platform with Inertial Measurements
arXiv:2512.19245v2 Announce Type: replace-cross Abstract: This paper tackles the problem of estimating the relative position, orientation, and velocity between …
SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation
arXiv:2605.12386v2 Announce Type: replace Abstract: Robotic manipulation is typically evaluated by task success, but successful completion does not guarantee sa…
Continual Quadruped Robots Coordination via Semantic Skill Discovery
arXiv:2606.08102v2 Announce Type: replace Abstract: Multi-quadruped coordination has attracted increasing attention due to its enhanced payload capacity, broade…
GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation
arXiv:2606.08530v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world de…
TORL-VLA: Tactile Guided Online Reinforcement Learning for Contact-Rich Manipulation
arXiv:2606.09337v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have become a powerful framework for robotic manipulation, and recent st…
Vision-Language-Action Jump-Starting for Reinforcement Learning Robotic Agents
arXiv:2604.13733v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) enables high-frequency, closed-loop control for robotic manipulation, but …
ActionMap: Robot Policy Learning via Voxel Action Heatmap
arXiv:2606.06904v2 Announce Type: replace Abstract: Vision-language-action (VLA) models have advanced rapidly across backbones, training recipes, and data scale…
Closing the Motion Execution Gap: From Semantic Motion Task Constraints to Kinematic Control
arXiv:2605.12053v2 Announce Type: replace Abstract: This paper addresses the Motion Execution Gap, the disconnect between high-level symbolic task descriptions …
EKF-Based Depth Camera and Deep Learning Fusion for UAV-Person Distance Estimation and Following in SAR Operations
arXiv:2602.20958v2 Announce Type: replace Abstract: Vision-based Unmanned Aerial Vehicles (UAVs) frameworks aid human search tasks by detecting and recognizing …
Phase-Based Multi-Gait Learning for a Salamander-Like Robot
arXiv:2511.08299v2 Announce Type: replace Abstract: Salamander-like robots are designed inspired by the skeletal structure of their biological counterparts. How…
The Unreasonable Effectiveness of Discrete-Time Gaussian Process Mixtures for Robot Policy Learning
arXiv:2505.03296v2 Announce Type: replace Abstract: We present Mixture of Discrete-time Gaussian Processes (MiDiGap), a novel approach for flexible policy repre…
Fast-SDE: Efficient Single-Microphone Sound Source Distance Estimation in Reverberant Environments
arXiv:2606.12339v1 Announce Type: cross Abstract: Sound source distance estimation (SDE) is a critical capability in human-robot interaction. An inappropriate i…
Making Foresight Actionable: Repurposing Representation Alignment in World Action Models
arXiv:2606.12217v1 Announce Type: cross Abstract: World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to …
DroneShield-AI: A Multi-Modal Sensor Fusion Framework for Real-Time Autonomous Drone Threat Detection, Behavioral Intent Classification, and Swarm Intelligence in Contested Airspace
arXiv:2606.11687v1 Announce Type: cross Abstract: Unmanned Aerial Vehicle (UAV) threats have emerged as a defining security challenge of the 21st century. This …
FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning
arXiv:2606.12406v1 Announce Type: new Abstract: Contact-rich manipulation requires force sensitivity, but many robot arms lack dedicated force sensors due to th…
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies
arXiv:2605.03065v2 Announce Type: replace-cross Abstract: Generative control policies (GCPs), such as diffusion- and flow-based control policies, have emerged a…
Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
arXiv:2605.00321v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) policies often fail under distribution shift, suggesting that decisions may dep…