Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesReliability-Guided RGB-D Sensor Fusion for Glare-Resilient Navigation Costmaps
arXiv:2604.12753v2 Announce Type: replace Abstract: Specular glare on reflective floors, glass boundaries, and glossy indoor surfaces can corrupt active-stereo …
Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving
arXiv:2609.04070v1 Announce Type: cross Abstract: Bridging the gap between the discrete reasoning of Vision-Language Models and the continuous, physics-constrai…
GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs
arXiv:2609.03892v1 Announce Type: cross Abstract: 3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in cu…
Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis
arXiv:2609.04096v1 Announce Type: new Abstract: This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizabl…
Toward Unified Robot Learning: Bridging Representation, Vision-Language-Action, and World Models
arXiv:2609.03927v1 Announce Type: new Abstract: For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and re…
A hybrid pipeline for dynamic ontology-based semantic mapping
arXiv:2609.03891v1 Announce Type: new Abstract: Semantic mapping plays a crucial role in the ability of a robot to interact with objects, operate and navigate a…
FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation
arXiv:2609.03889v1 Announce Type: new Abstract: Contact-rich loco-manipulation requires a bridge between semantic action generation and physical interaction con…
A Multi-Vine Soft Robot Enabling Accessible Working Channel and Steering
arXiv:2609.03758v1 Announce Type: new Abstract: Soft eversion robots, also known as vine robots, have attracted growing interest for navigation and inspection t…
MINERVA: How Small Can a Manipulation Policy Be and Still Solve LIBERO?
arXiv:2609.03715v1 Announce Type: new Abstract: Vision-language-action (VLA) models with billions of parameters now dominate the LIBERO manipulation benchmark, …
Local Path Planning and Obstacle Avoidance for an Omnicopter Platform
arXiv:2609.03630v1 Announce Type: new Abstract: Autonomous unmanned aerial vehicles (UAVs) increasingly operate in cluttered environments where global planners …
Establishing a Dynamic Multimodal HRI Dataset for Engagement Analysis with a Humanoid Robot
arXiv:2609.03255v1 Announce Type: new Abstract: This paper presents an experimental design for constructing a multimodal dataset to analyze user engagement in h…
Long-Horizon Consistent and Interaction-Aware World Models for Multi-Style End-to-End Driving
arXiv:2609.03225v1 Announce Type: new Abstract: End-to-end autonomous driving has increasingly adopted world model-based reinforcement learning frameworks to im…
Sensing Which Modality Matters: Evidence-Gated Regularization for Robust VLA Policies
arXiv:2609.03142v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies fuse multimodal sensory inputs, but training on limited and homogeneous ro…
GPU-Accelerated Astrodynamics World Models for Spacecraft Rendezvous and Proximity Operations
arXiv:2609.03067v1 Announce Type: new Abstract: World models are an emerging paradigm in representation learning in which an agent jointly learns state-action d…
Real-Time Shape Control of Multi-Segment Soft Robotic Arms Using Koopman Operators with Global and Local Observables
arXiv:2609.03175v1 Announce Type: new Abstract: Multi-segment soft robotic arms can continuously reconfigure their body shapes for safe interaction, but tip con…
R2S-Eval: Robot Evaluation with Real-to-Sim Calibration via Vision-Language Models
arXiv:2609.03276v1 Announce Type: new Abstract: Evaluating robot manipulation policies is becoming increasingly important as generalist models, particularly vis…
WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models
arXiv:2609.03681v1 Announce Type: new Abstract: Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinfor…
GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation
arXiv:2609.04193v1 Announce Type: new Abstract: Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic…
One Demonstration, Many Objects: Generalizing Manipulation via Local Contact Geometry
arXiv:2609.01938v2 Announce Type: replace Abstract: Dexterous manipulation with multi-fingered robot hands promises human-level dexterity, but collecting large-…
Towards Lifelong Aerial Autonomy: Geometric Memory Management for Continual Visual Place Recognition in Dynamic Environments
arXiv:2604.09038v2 Announce Type: replace Abstract: Robust geo-localization under changing environmental and operational conditions is critical for long-term ae…
Highly Deformable Proprioceptive Membrane for Real-Time 3D Shape Reconstruction
arXiv:2601.13574v3 Announce Type: replace Abstract: Reconstructing the three-dimensional (3D) geometry of object surfaces is essential for robot perception, yet…
A Quantitative Comparison of Centralised and Distributed Reinforcement Learning-Based Control for Soft Robotic Arms
arXiv:2511.02192v3 Announce Type: replace Abstract: This paper presents a quantitative comparison between centralised and distributed multi-agent reinforcement …
Vision-Based Tactile Sensing for the Perception of the Object's Compliance and Hardness
arXiv:2510.12528v2 Announce Type: replace Abstract: Object compliance perception enables the identification of soft materials, supporting tasks such as fruit de…
LightEMMA: A Longitudinal Evaluation of Vision-Language Models for Autonomous Driving
arXiv:2505.00284v3 Announce Type: replace Abstract: Rapid advances in vision-language models (VLMs) have generated growing interest in their application to auto…
DogLegs: Robust Proprioceptive State Estimation for Legged Robots Using Multiple Leg-Mounted IMUs
arXiv:2503.04580v3 Announce Type: replace Abstract: Robust and accurate proprioceptive state estimation of the main body is crucial for legged robots to execute…
Dancing with REEM-C: A robot-to-human physical-social communication study
arXiv:2408.05301v3 Announce Type: replace Abstract: Humans often work closely together and relay a wealth of information through physical interaction. Robots, o…
Subspace Inference Enables Efficient Active Reward Learning from Preferences
arXiv:2609.04066v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has emerged as a powerful yet sample-inefficient approach fo…
Rethinking World Models for Safety-Critical Embodied Systems
arXiv:2609.03774v1 Announce Type: cross Abstract: World models have progressed from compact latent dynamics to generative, controllable, and interactive simulat…
RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning
arXiv:2609.03199v1 Announce Type: cross Abstract: Robot learning increasingly depends on broad and diverse demonstrations, yet collecting robot data remains exp…
SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving
arXiv:2609.03602v1 Announce Type: cross Abstract: World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive…