Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesReSiReg: Towards Spatially Consistent Semantics in Language-Conditioned Robotic Tasks
arXiv:2606.19088v1 Announce Type: new Abstract: Vision-Language Models (VLMs) enable robots to follow open-language instructions. However, dense VLM embeddings …
GCNGrasp-VP: Affordance-Guided View Planning for Efficient Task-Oriented Grasping
arXiv:2606.19091v1 Announce Type: new Abstract: Task-oriented grasping performance degrades significantly when object views suffer from occlusions. Existing tas…
DexSynRefine: Synthesizing and Refining Human-Object Interaction Motion for Physically Feasible Dexterous Robot Actions
arXiv:2605.05925v2 Announce Type: replace Abstract: Learning dexterous manipulation from human-object interaction (HOI) data offers a scalable alternative to ro…
Viking Hill Dataset: A Lidar-Radar-Camera Dataset for Detection and Segmentation in Forest Scenes
arXiv:2606.19154v1 Announce Type: new Abstract: Autonomous robots operating under forest canopies need robust perception of trees and surrounding vegetation acr…
Hardware- and Vision-in-the-Loop Validation of Deep Monocular Pose Estimation for Autonomous Maritime UAV Flight
arXiv:2606.19176v1 Announce Type: new Abstract: Autonomous UAV operations on ships require reliable vision-based relative pose estimation, yet at-sea validation…
Seeing Through Occlusion: Deterministic Arm Kinematic Correction for Robot Teleoperation
arXiv:2606.19240v1 Announce Type: new Abstract: Markerless, single-RGB-D-camera motion capture provides a low-cost and non-invasive alternative to conventional …
Observability and Consistency Analysis for Visual-Inertial Navigation with Anchored Feature Parameterizations
arXiv:2606.19307v1 Announce Type: new Abstract: This paper presents an analysis of the observability and consistency properties of filtering-based visual-inerti…
Modeling Branches for Active Manipulation using Iterative Parameter Estimation
arXiv:2606.19314v1 Announce Type: new Abstract: This study presents a method for modeling diverse plant branches by iteratively estimating material parameters t…
Do as I Do: Dexterous Manipulation Data from Everyday Human Videos
arXiv:2606.19333v1 Announce Type: new Abstract: How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous…
Technical Report for ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Leveraging DINOv3 for Robust Outdoor Scene Understanding in Field Robotics
arXiv:2606.18582v1 Announce Type: cross Abstract: The GOOSE 2D Fine-Grained Semantic Segmentation Challenge at the ICRA 2026 Workshop on Field Robotics evaluate…
Mutual Adaptation in Human-Robot Co-Transportation with Human Preference Uncertainty
arXiv:2503.08895v2 Announce Type: replace Abstract: Mutual adaptation can enhance overall task performance in human-robot co-transportation by integrating both …
STORM: Slot-based Task-aware Object-centric Representation for robotic Manipulation
arXiv:2601.20381v2 Announce Type: replace Abstract: Visual foundation models provide strong perceptual features for robotics, but their dense representations la…
Cosmos 3: Omnimodal World Models for Physical AI
arXiv:2606.02800v3 Announce Type: replace-cross Abstract: We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate lan…
Robust and Efficient MuJoCo-based Model Predictive Control via Web of Affine Spaces Derivatives
arXiv:2512.21109v2 Announce Type: replace Abstract: MuJoCo is a powerful and efficient physics simulator widely used in robotics. One common way it is applied i…
Odyssey: An Automotive Lidar-Inertial Odometry Dataset with GNSS-denied situations
arXiv:2512.14428v2 Announce Type: replace Abstract: The development and evaluation of Lidar-Inertial Odometry (LIO) and Simultaneous Localization and Mapping (S…
Enhancing Fatigue Detection through Heterogeneous Multi-Source Data Integration and Cross-Domain Modality Imputation
arXiv:2507.16859v5 Announce Type: replace Abstract: Fatigue detection for human operators is important in safety-related applications such as aviation, mining, …
UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning
arXiv:2606.19328v1 Announce Type: cross Abstract: Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, byp…
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models
arXiv:2606.19297v1 Announce Type: cross Abstract: Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on…
N(CO)$^2$: Neural Combinatorial Optimization with Chance Constraints to Solve Stochastic Orienteering
arXiv:2606.18514v1 Announce Type: new Abstract: Neural combinatorial optimization (NCO) offers a promising alternative to traditional heuristic-based methods fo…
Task Allocation and Motion Planning in Dynamic, Cluttered Environments via CBBA and Graphs of Convex Sets
arXiv:2606.18516v1 Announce Type: new Abstract: Multi-agent task planning in cluttered, dynamic environments requires assigning tasks to agents while simultaneo…
Selective Unit-Cell Actuation in Lattice Structures for Distributed Morphology in Soft Robots
arXiv:2606.18704v1 Announce Type: new Abstract: Soft lattice structures are increasingly used in robotics to tailor compliance and guide deformation; however, a…
RSLCPP -- Deterministic Simulations Using ROS 2
arXiv:2601.07052v2 Announce Type: replace Abstract: Simulation is crucial in real-world robotics, offering safe, scalable, and efficient environments for develo…
CABLE: Cloud-Assisted Bandwidth-efficient LMM-based Encoding for V2X Systems
arXiv:2606.19258v1 Announce Type: cross Abstract: Cloud-hosted large multimodal models (LMMs) can provide strong open-vocabulary perception for Vehicle-to-Every…
Aerial-ground LiDAR place recognition with patch-level self-supervised learning and expanded reciprocal re-ranking
arXiv:2606.18583v1 Announce Type: cross Abstract: LiDAR place recognition determines one's position on a prior point cloud map. The most studied ground-level Li…
Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos
arXiv:2606.18955v1 Announce Type: cross Abstract: Training generalist Vision-Language-Action(VLA) models typically requires massive, diverse robotic datasets wi…
Stealthy World Model Manipulation via Data Poisoning
arXiv:2606.18697v1 Announce Type: cross Abstract: Model-based learning agents use learned world models to predict future states, plan actions, and adapt to new …
AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework
arXiv:2606.18532v1 Announce Type: cross Abstract: AI systems are increasingly evaluated in bounded environments that combine isolation, simulation, instrumentat…
RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer
arXiv:2606.18439v1 Announce Type: cross Abstract: Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one fo…
A Mixed-Reality Testbed for Autonomous Vehicles
arXiv:2606.19267v1 Announce Type: new Abstract: We propose a mixed-reality, hardware-in-the-loop (HIL) testbed for autonomous vehicles that seamlessly integrate…
Constant Time-Delay Leader Following with Neural Networks and Invariant Extended Kalman Filters for Arbitrary Trajectories
arXiv:2606.19227v1 Announce Type: new Abstract: This paper proposes a constant time-delay trajectory tracking method for vehicle convoys operating without inter…