Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesFlexLAM: Resolving the Bottleneck Trade-off in Latent Action Learning
arXiv:2606.19408v1 Announce Type: cross Abstract: Latent actions provide a compact interface between action-free video and downstream decision-making, yet exist…
CTS-MoE: Implicit Terrain Adaptation via Mixture-of-Experts for Perceptive Locomotion
arXiv:2606.19633v1 Announce Type: new Abstract: Perceptive legged locomotion over discontinuous terrain (e.g., stairs, gaps, and obstacles) requires adaptive be…
Fail-RAG : A Retrieval Augmented Generation Informed Framework for Robot Failure Identification
arXiv:2606.19598v1 Announce Type: new Abstract: Industry automation is witnessing an evolution in robotics driven by both technological breakthroughs and societ…
pdSTL: Probabilistic Differentiable Signal Temporal Logic for Stochastic Systems
arXiv:2606.19561v1 Announce Type: new Abstract: Autonomous robots operating in uncertain environments must satisfy complex temporal and safety specifications de…
Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning
arXiv:2606.19340v1 Announce Type: new Abstract: We present a zero-shot framework for long-horizon dexterous manipulation that grounds language instructions into…
Spatially Stratified Distillation for Heterogeneous Radar Place Recognition
arXiv:2606.18687v1 Announce Type: cross Abstract: Scalable, all-weather place recognition increasingly relies on heterogeneous radar place recognition to bridge…
Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation
arXiv:2606.18960v1 Announce Type: cross Abstract: Action-conditioned world models have emerged as a promising paradigm for robot learning, offering a scalable a…
OneCanvas: 3D Scene Understanding via Panoramic Reprojection
arXiv:2606.19253v1 Announce Type: cross Abstract: Existing approaches to 3D scene understanding in Vision-Language Models (VLMs) either rely on complex, model-s…
Steering Flexible Linear Objects in Planar Environments by Two Robot Hands Using Euler's Elastica Solutions
arXiv:2501.02874v5 Announce Type: replace Abstract: The manipulation of flexible objects such as cables, wires and fresh food items by robot hands forms a speci…
R2BC: Multi-Agent Imitation Learning from Single-Agent Demonstrations
arXiv:2510.18085v2 Announce Type: replace Abstract: Imitation Learning (IL) is a natural way for humans to teach robots, particularly when high-quality demonstr…
TurboMap: GPU-Accelerated Local Mapping for Visual SLAM
arXiv:2511.02036v4 Announce Type: replace Abstract: In real-time Visual SLAM systems, local mapping must operate under strict latency constraints, as delays deg…
Bench-Push: Benchmarking Pushing-based Navigation and Manipulation Tasks for Mobile Robots
arXiv:2512.11736v2 Announce Type: replace Abstract: Mobile robots are increasingly deployed in cluttered environments with movable objects, posing challenges fo…
Benchmarking Action Spaces in Reinforcement Learning for Vision-based Robotic Manipulation
arXiv:2606.18594v1 Announce Type: new Abstract: In real-world reinforcement learning (RL), the choice of action space can play a key role in shaping motion smoo…
Tilt-Ropter: A Fully Actuated Hybrid Aerial-Terrestrial Vehicle with Tilt Rotors and Passive Wheels
arXiv:2602.01700v2 Announce Type: replace Abstract: In this work, we present Tilt-Ropter, a fully actuated hybrid aerial-terrestrial vehicle (HATV) that integra…
Quantile Transfer for Reliable Operating Point Selection in Visual Place Recognition
arXiv:2602.04401v2 Announce Type: replace Abstract: Visual Place Recognition (VPR) is a key component for localisation in Global Navigation Satellite System (GN…
Embedding Semantic Risk into Distance Fields and CBFs for Online Monocular Safe Control
arXiv:2606.01605v2 Announce Type: replace Abstract: We propose an online monocular perception-to-control framework that embeds semantic risk into the distance f…
SRL: Combining SLIP Model and Reinforcement Learning for Agile Robotic Jumping
arXiv:2606.18625v1 Announce Type: new Abstract: Robotic jumping is pivotal in applications such as search and rescue and logistics, where crossing obstacles and…
VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision
arXiv:2606.18426v1 Announce Type: new Abstract: We introduce VEGA, an approach for training navigation VisionLanguage-Action (VLA) models from unlabeled egocent…
High-Degree-of-Freedom Lightweight Bioinspired Leg for Enhanced Mobility in Small Robots
arXiv:2606.18680v1 Announce Type: new Abstract: In microrobotics, enhancing locomotion capabilities by increasing the degrees of freedom (DoF) of leg mechanisms…
PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation
arXiv:2606.18375v1 Announce Type: new Abstract: World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting …
Why Automate This? Exploring Correlations Between Desire for Robotic Automation, Invested Time and Well-Being
arXiv:2501.06348v4 Announce Type: replace-cross Abstract: Understanding the motivations underlying the human inclination to automate tasks is vital for developi…
Guava: An Effective and Universal Harness for Embodied Manipulation
arXiv:2606.18363v1 Announce Type: new Abstract: Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agen…
Recover, Discover, Plan: Learning Skills and Concepts from Robot Failures
arXiv:2606.18328v1 Announce Type: new Abstract: Intelligent robots should not only recover from failures, but also acquire the abstract knowledge needed to avoi…
Congestion-Aware Robot Tour Planning in Crowded Environments
arXiv:2606.19031v1 Announce Type: new Abstract: Autonomous mobile service robots are often required to complete tours that require navigating through a set of l…
ReSiReg: Towards Spatially Consistent Semantics in Language-Conditioned Robotic Tasks
arXiv:2606.19088v1 Announce Type: new Abstract: Vision-Language Models (VLMs) enable robots to follow open-language instructions. However, dense VLM embeddings …
GCNGrasp-VP: Affordance-Guided View Planning for Efficient Task-Oriented Grasping
arXiv:2606.19091v1 Announce Type: new Abstract: Task-oriented grasping performance degrades significantly when object views suffer from occlusions. Existing tas…
DexSynRefine: Synthesizing and Refining Human-Object Interaction Motion for Physically Feasible Dexterous Robot Actions
arXiv:2605.05925v2 Announce Type: replace Abstract: Learning dexterous manipulation from human-object interaction (HOI) data offers a scalable alternative to ro…
Viking Hill Dataset: A Lidar-Radar-Camera Dataset for Detection and Segmentation in Forest Scenes
arXiv:2606.19154v1 Announce Type: new Abstract: Autonomous robots operating under forest canopies need robust perception of trees and surrounding vegetation acr…
Hardware- and Vision-in-the-Loop Validation of Deep Monocular Pose Estimation for Autonomous Maritime UAV Flight
arXiv:2606.19176v1 Announce Type: new Abstract: Autonomous UAV operations on ships require reliable vision-based relative pose estimation, yet at-sea validation…
Seeing Through Occlusion: Deterministic Arm Kinematic Correction for Robot Teleoperation
arXiv:2606.19240v1 Announce Type: new Abstract: Markerless, single-RGB-D-camera motion capture provides a low-cost and non-invasive alternative to conventional …