Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesDevelopment of a 3 in Sewer Pipe Inspection Robot with an Articulated Differential Mechanism using X-shaped Linkages
arXiv:2606.14070v1 Announce Type: new Abstract: This paper proposes, an improved version of the 3 in sewer pipe inspection robot equipped with an emergency evac…
WAM4D: Fast 4D World Action Model via Spatial Register Tokens
arXiv:2606.14048v1 Announce Type: cross Abstract: World action models (WAMs) have recently shown promise in jointly modeling future observations and executable …
CADET: Physics-Grounded Causal Auditing and Training-Free Deconfounding of End-to-End Driving Planners
arXiv:2606.14438v1 Announce Type: new Abstract: End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they assoc…
Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis
arXiv:2606.08881v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet exi…
Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
arXiv:2606.14409v1 Announce Type: new Abstract: In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the fu…
Short-Horizon Position Accuracy of Single-Track Models: Implications for Motion Planning of Autonomous Vehicles
arXiv:2606.14216v1 Announce Type: new Abstract: Accurate and computationally efficient vehicle models are essential for motion planning of autonomous vehicles, …
Elastic Queries Reinforcement Learning: Self-Aware Policy Execution for VLA Models
arXiv:2606.14375v1 Announce Type: new Abstract: Vision-language-action (VLA) models are powerful action generators for robot manipulation, but they are typicall…
ForestBack: Breadcrumb-Based Pedestrian Dead Reckoning for Infrastructure-Free Return Navigation
arXiv:2606.14421v1 Announce Type: new Abstract: Reliable return navigation remains an important challenge in GPS-denied environments where external positioning …
From Attacks to Curricula: Learnability-Guided Adversarial Training for Safe Autonomous Driving
arXiv:2606.14032v1 Announce Type: new Abstract: Closed-loop adversarial training improves autonomous driving safety by exposing policies to rare safety-critical…
SplatlessDF: Continuous Distance Field Mapping with Non-Splatting Gaussians
arXiv:2606.13990v1 Announce Type: new Abstract: Recent Gaussian splatting (GS) methods have shown that scenes can be represented efficiently with optimisable Ga…
ORCA: A Platform for Open-Source Dexterity Research
arXiv:2606.14561v1 Announce Type: new Abstract: Robotics manipulation research increasingly focuses on two-finger parallel grippers for their effectiveness, aff…
AnyGoal: Vision-Language Guided Multi-Agent Exploration for Training-Free Lifelong Navigation
arXiv:2606.13878v1 Announce Type: new Abstract: End-to-end navigation policies trained on large simulation corpora degrade sharply when transferred to out-of-di…
Scalable Dynamic Tactile Sensing Enabled by Passive and Flexible Acoustic Waveguides
arXiv:2606.13746v1 Announce Type: new Abstract: Artificial dynamic tactile sensing requires sensitivity, robustness, and compliance, yet existing technologies f…
EWAM: An Enhanced World Action Model for Closed-Loop Online Adaptation in Embodied Intelligence
arXiv:2606.12690v1 Announce Type: new Abstract: In this paper, we propose the Enhanced World Action Model (EWAM), a closed-loop online adaptation architecture b…
EquiDexFlow: Contact-Grounded SE(3)-Equivariant Dexterous Grasp Generative Flows
arXiv:2606.12728v1 Announce Type: new Abstract: Most learned dexterous grasp generators relegate contact forces to a downstream verification step, so a kinemati…
Sparse2Act: Learning Action-Aligned Sparse 3D Representations for Cross-Domain Robot Manipulation
arXiv:2606.12759v1 Announce Type: new Abstract: Explicit 3D representations are attractive for manipulation because they expose object shape, workspace geometry…
EmbodiSteer: Steering Embodiment-Agnostic Visuomotor Policies with Joint-Space Guidance for Zero-Shot Cross-Embodiment Deployment
arXiv:2606.12965v1 Announce Type: new Abstract: Scalable robot imitation learning relies on large-scale heterogeneous data from diverse robots or body-free data…
Y-BotFrame: An Extensible Embodied Agent Framework for Quadruped Robot Assistants
arXiv:2606.13049v1 Announce Type: new Abstract: Quadruped robots are capable of traversing a wide range of complex terrains with high flexibility. As highly mob…
EA-WM: Event-Aware World Models with Task-Specification Grounding for Long-Horizon Manipulation
arXiv:2606.13053v1 Announce Type: new Abstract: Pretrained-feature world models provide a useful substrate for robot imagination, but visual or latent predictio…
Redesigning Regularization for Effective Policy Smoothing
arXiv:2606.13169v1 Announce Type: new Abstract: This paper proposes a novel regularization design to effectively smooth policy functions in reinforcement learni…
WT-UMI: Tactile-based Whole-Body Manipulation via Force-Supervised Contact-Aware Planning
arXiv:2606.13232v1 Announce Type: new Abstract: Whole-body humanoid manipulation of bulky, deformable, and shared-load objects requires distributed contact sens…
MCR-Bionic Hand: Anatomical Structural Priors for Dexterous Manipulation
arXiv:2606.13601v1 Announce Type: new Abstract: Dexterous robotic hands are usually formulated as high dimensional active control systems governed by degrees of…
Scale Buys Interpolation, Structure Buys a Horizon: Certified Predictability for Equivariant World Models
arXiv:2606.13092v1 Announce Type: cross Abstract: Scale buys interpolation; structure buys a certified horizon. A world model's average error says nothing about…
MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models
arXiv:2606.13515v1 Announce Type: cross Abstract: World Action Models (WAMs) present a promising paradigm for robotic control via video prediction. However, cur…
LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories
arXiv:2606.13578v1 Announce Type: cross Abstract: Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of d…
Learning Robot Safety from Sparse Human Feedback using Conformal Prediction
arXiv:2501.04823v2 Announce Type: replace Abstract: Ensuring robot safety can be challenging; user-defined constraints can miss edge cases, policies can become …
$\texttt{WEAVER}$, Better, Faster, Longer: An Effective World Model for Robotic Manipulation
arXiv:2606.13672v1 Announce Type: new Abstract: The potential impacts of world models (WMs, i.e., learned simulators) on robotics are far-reaching -- policy eva…
NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation
arXiv:2606.13494v1 Announce Type: new Abstract: Goal-conditioned visual navigation requires a robot to act under partial observability by anticipating how its m…
Humor Style Drives Laughter, Topic Shapes Acceptability: Evaluating Bilingual Personal and Political Robot-Delivered AI Jokes
arXiv:2606.13256v1 Announce Type: new Abstract: Humor plays a central role in human social relationships, and recent advances in computational humor create new …
Embedding ISO 10218 Safety Compliance in Robots via Control Barrier Functions for Human-Robot Collaboration
arXiv:2606.13203v1 Announce Type: new Abstract: Human-Robot Collaboration (HRC) requires strict adherence to safety standards, such as ISO 10218, to prevent har…