Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesStellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models
arXiv:2608.11671v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often …
Policy-Induced Hand Priors in Humanoid Dual-Arm Manipulation: Diagnosing and Mitigating Initial-Pose Dependence
arXiv:2608.11769v1 Announce Type: new Abstract: Vision-language-action (VLA) policies are expected to operate robustly across variations in the robot's initial …
Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)
arXiv:2608.11229v1 Announce Type: cross Abstract: Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align rob…
Self-Evolving Embodied Agents via Skill-Harness Evolution
arXiv:2608.11350v1 Announce Type: cross Abstract: Embodied agents are increasingly built as systems around foundation models, where performance depends not only…
Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards
arXiv:2608.11451v1 Announce Type: new Abstract: Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that…
World Tokens: Enhancing Embodied Policies with Training-Time World Modeling
arXiv:2608.09730v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficie…
Scalable Multi-Agent Maze Traversal with Local Communication
arXiv:2608.11895v1 Announce Type: new Abstract: Cave networks, pipe systems, and similar maze-like environments pose significant challenges for multi-agent navi…
Enhancing Visual Domain Robustness in Behaviour Cloning via Saliency-Guided Augmentation
arXiv:2608.11870v1 Announce Type: new Abstract: In vision-based behavior cloning (BC), conventional image augmentations such as Random Crop and Color Jitter oft…
G0.5: One Autoregressive Stream for Robot Reasoning and Action
arXiv:2608.11739v1 Announce Type: new Abstract: The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained…
Adaptation of Generalist Robot Policies with Minimal Data
arXiv:2608.11363v1 Announce Type: new Abstract: A central goal in robot learning is to move beyond task-specific human data collection toward robots that improv…
Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment
arXiv:2608.12198v1 Announce Type: new Abstract: Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches…
HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing
arXiv:2608.12122v1 Announce Type: new Abstract: Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the hi…
Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
arXiv:2608.12063v1 Announce Type: new Abstract: Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Lear…
DaViNCi: A Dataset Towards Outdoor Vision-and-Language Navigation with Continuous Actions and Dynamic Elements
arXiv:2608.11901v1 Announce Type: new Abstract: Vision-and-Language Navigation (VLN) has progressively expanded from indoor to outdoor environments. However, ex…
D3D-GEN: Robot-Aware Domain-Grounded Interactive 3D World Generation for Social Robotics
arXiv:2608.11876v1 Announce Type: new Abstract: Training and validation of Embodied AI for social navigation critically depends on realistic simulation environm…
Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds
arXiv:2608.10056v1 Announce Type: new Abstract: Following a target human in crowded environments involves an inherent conflict between staying close to the targ…
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
arXiv:2603.19201v3 Announce Type: replace Abstract: Contact-rich manipulation tasks, such as wiping and assembly, require accurate perception of contact forces,…
Navigating in Uncertain Environments with Heterogeneous Visibility
arXiv:2603.03495v2 Announce Type: replace Abstract: Navigating an environment with uncertain connectivity requires a strategic balance between minimizing the co…
Bandwidth-Efficient Multi-Agent Communication through Information Bottleneck and Vector Quantization
arXiv:2602.02035v2 Announce Type: replace Abstract: Multi-agent reinforcement learning systems deployed in real-world robotics applications face severe communic…
Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)
arXiv:2509.06191v2 Announce Type: replace Abstract: Recent 3D generative models, which are capable of generating full object shapes from just a few images, now …
GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes
arXiv:2608.10886v1 Announce Type: cross Abstract: Robots operating in human environments need memories that capture not only what objects exist and where, but a…
Robust Safety Filtering for Input-Constrained Underactuated Linear Systems
arXiv:2608.10872v1 Announce Type: cross Abstract: We present a robust safety-filtering framework for input-constrained underactuated linear systems subject to u…
The Impact of Operational-Data Fidelity when Assessing Safety-Critical Autonomous-Vehicle Software
arXiv:2608.10025v1 Announce Type: new Abstract: For safety-critical software, data from the software's operational past (e.g. a sequence of success and failure …
Protection Levels for Vision-Based Pose Estimation
arXiv:2608.10023v1 Announce Type: new Abstract: Vision-based navigation complements Global Navigation Satellite Systems, but certification demands integrity gua…
AECNav: Active Evidence Consolidation for Efficient Zero-Shot Open-Vocabulary Object Navigation
arXiv:2608.10817v1 Announce Type: new Abstract: Zero-shot object-goal navigation (ZSON) in open-vocabulary scenarios is challenging, as it requires a robot to l…
Enabling Scalable Kinesthetic Teaching via Observer-based Hand-guiding with Active Support
arXiv:2608.10847v1 Announce Type: new Abstract: Kinesthetic teaching through robot hand-guiding provides a natural interface for collecting demonstrations in im…
Aerial Layouting: Design and Control of a Compliant and Actuated End-Effector for Precise In-flight Marking on Ceilings
arXiv:2608.10987v1 Announce Type: new Abstract: Aerial robots have demonstrated impressive feats of precise control, such as dynamic flight through openings or …
VIScore: Diagnosing Planning-Relevant Quality in Latent World Models
arXiv:2608.11174v1 Announce Type: new Abstract: Regulating the latent space to an isotropic Gaussian distribution provides a stable and information-maximized la…
Risk-Aware Kinodynamic Motion Planning Under Uncertainty For Safe Navigation on Planetary Environments
arXiv:2608.11175v1 Announce Type: new Abstract: For autonomous space exploration, robotic agents need to perform motion planning in which environmental interact…
Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning
arXiv:2608.11204v1 Announce Type: new Abstract: Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstration…