Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesEmbodied GPT-5.1: Evidence of a World Model?
arXiv:2607.23899v1 Announce Type: new Abstract: This exploratory study examines whether a large multimodal language model, GPT-5.1, can serve as the high-level …
WorldDiT: A Unified Diffusion Architecture for World and Action Modeling
arXiv:2607.23909v1 Announce Type: cross Abstract: Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as …
Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models
arXiv:2607.23602v1 Announce Type: new Abstract: Controllers based on sampling and latent world models assign a predicted terminal cost to each candidate action …
NEO: NeRF It Once, Edit It Many Times for Continuous Object Manipulation
arXiv:2607.24538v1 Announce Type: new Abstract: In this paper, we present NEO, a unified framework providing language-guided NeRF editing for robotic manipulati…
ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm
arXiv:2607.24481v1 Announce Type: new Abstract: Real-world evaluation is a bottleneck in developing generalist robot manipulation policies. Each rollout require…
LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratories
arXiv:2607.23704v1 Announce Type: new Abstract: The deployment of embodied agents in self-driving laboratories could accelerate scientific discovery, yet their …
Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline
arXiv:2607.22997v1 Announce Type: new Abstract: Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the…
SimBEV2X: A Large-Scale Dataset and Data Generation Tool for Multi-Task Vehicle-to-Everything Cooperative Perception
arXiv:2607.23910v1 Announce Type: cross Abstract: Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical li…
Kernel-SDF: An Open-Source Library for Real-Time Signed Distance Function Estimation using Kernel Regression
arXiv:2603.29227v2 Announce Type: replace Abstract: Accurate and efficient scene representation is crucial for robotic tasks such as motion planning, manipulati…
Observer-Assisted Relative-Velocity Compensation with LPV-$H_\infty$ Robust Correction for 3D Trajectory Tracking of Underactuated Non-Minimum-Phase AUVs under Ocean Currents
arXiv:2607.23653v1 Announce Type: cross Abstract: This paper develops an observer-assisted control architecture for 3D trajectory tracking of torpedo-type under…
Anticipatory Risk-Guided Reinforcement Learning for Safe Flight Through Dynamic Clutter
arXiv:2607.23565v1 Announce Type: new Abstract: Safe quadrotor navigation in cluttered and dynamic environments depends not only on instantaneous geometric perc…
Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot
arXiv:2607.23204v1 Announce Type: new Abstract: Large language model (LLM)-based dialogue systems suffer response delays because generation begins only after fi…
Data Pyramid for Embodied Manipulation
arXiv:2607.24744v1 Announce Type: new Abstract: Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit …
A note on the motion representation and configuration update in time stepping schemes for the constrained rigid body
arXiv:2607.24317v1 Announce Type: new Abstract: The dynamics of a holonomically constrained rigid body can be modeled by Newton-Euler equations subjected to geo…
Learning Adaptive Multi-Task Guidance, Navigation, and Control via Hypernetworks
arXiv:2607.24292v1 Announce Type: new Abstract: Autonomous free-flying robots in orbital environments require controllers that are both versatile and resource-e…
Phenology-based learning framework for yield estimation and harvest forecasting of raspberry fruits
arXiv:2411.00967v2 Announce Type: replace-cross Abstract: The future of agriculture is intertwined with automation. Accurate fruit detection, yield estimation, …
LAGS: Low-Altitude Gaussian Splatting with Groupwise Heterogeneous Graph Learning
arXiv:2604.16910v2 Announce Type: replace-cross Abstract: Low-altitude Gaussian splatting (LAGS) facilitates 3D scene reconstruction by aggregating aerial image…
Distributed Coordination for Resilient Multi-UAV Remote Sensing: A Photovoltaic Inspection Case Study
arXiv:2607.24482v1 Announce Type: new Abstract: Deploying multiple UAVs for remote sensing enables proportional reductions in mission time, but realizing these …
MoE-Based Learned Inertial Odometry for Bicycle Localization
arXiv:2510.17604v2 Announce Type: replace Abstract: GNSS suffers from multipath errors in urban canyons, making reliable bicycle localization difficult. Hand-cr…
Development of a Handheld Actuation Mechanism for a Tendon-driven Robotically Steered Guidewire
arXiv:2607.24629v1 Announce Type: new Abstract: An endovascular intervention begins with a skilled clinician manually navigating a long, slender wire, called a …
Quality-Adaptive Multi-UAV 3D Reconstruction with Sparse Workload Redistribution
arXiv:2607.24233v1 Announce Type: new Abstract: 3D reconstruction of unknown environments is a key application in robotics but is severely limited by the comput…
Learning Traversability-Aware Global Planners for Long Horizon Off-Road Navigation
arXiv:2607.23743v1 Announce Type: new Abstract: Autonomous navigation across large off-road environments remains a challenging problem. Onboard sensors perceive…
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
arXiv:2512.24125v3 Announce Type: replace Abstract: General-purpose robotic systems operating in open-world environments must achieve both broad generalization …
{\tau}: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
arXiv:2607.24485v1 Announce Type: new Abstract: Learning the informative tactile representation while effectively adapting it to pretrained Vision-Language-Acti…
MAGS-SLAM: Monocular Multi-Agent Gaussian Splatting SLAM for Geometrically and Photometrically Consistent Reconstruction
arXiv:2605.10760v2 Announce Type: replace Abstract: Collaborative photorealistic 3D reconstruction from multiple agents enables rapid large-scale scene capture …
Hierarchical Soft Actor-Critic for Sparse-Reward Long-Horizon Reinforcement Learning
arXiv:2607.23726v1 Announce Type: new Abstract: Exploration in sparse-reward long-horizon tasks poses significant challenges for reinforcement learning. To addr…
Memory for Attention: Language-Conditioned Re-Perception with a Vision--Language--Motion Map
arXiv:2607.23797v1 Announce Type: new Abstract: A robot carrying a persistent, behavior-annotated map faces two planning questions, and its memory answers only …
Cost-Aware Recovery-Pathway Identification and Bayesian Optimization for Autonomous Materials Discovery
arXiv:2607.23896v1 Announce Type: cross Abstract: Autonomous laboratories automate experimental execution, but a campaign must also decide which recovery pathwa…
Physical AI Governance: From Theory to Practice Across Life Cycle
arXiv:2607.22877v1 Announce Type: cross Abstract: With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to em…
SHARE: Towards Head-Mounted AR with User-Centric SLAM in Shared Human-Robot Workspaces
arXiv:2607.23901v1 Announce Type: cross Abstract: Human-Robot Collaboration (HRC) in shared physical spaces using Augmented Reality (AR) interfaces is powered b…