Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2826 storiesIndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
arXiv:2511.17384v2 Announce Type: replace Abstract: While Visual Large Language Models (VLLMs) show great promise as embodied agents, they continue to face subs…
Humanoid says KinetIQ Ascend reinforcement learning approaches human-level dexterity
Humanoid says its KinetIQ Ascend approach can reach 99.9% manipulation reliability at human speed and beyond for industrial tasks. The post Humanoid says KinetI…
Context is king: How Avride uses cloud VLMs as a safety net for delivery robots
Avride uses vision-language models, or VLMs, to improve the environmental awareness of its delivery robots. The post Context is king: How Avride uses cloud VLMs…
Choreographing the Way of Water: A Computational Framework for Aquatic Robotic Art
arXiv:2607.02174v1 Announce Type: new Abstract: Robotic choreography in open water is governed by nonlinear fluid dynamics, which impose significant challenges …
Learning to Localize Reference Trajectories in Image-Space for Visual Navigation
arXiv:2602.18803v2 Announce Type: replace Abstract: We present LoTIS, a model for visual navigation that provides robot-agnostic image-space guidance by localiz…
BIEVR-LIO: Robust LiDAR-Inertial Odometry through Bump-Image-Enhanced Voxel Maps
arXiv:2604.14421v2 Announce Type: replace Abstract: Reliable odometry is essential for mobile robots as they increasingly enter more challenging environments, w…
Simulation Based Reward Function Validation for Multi-Agent On Orbit Inspection
arXiv:2607.01367v1 Announce Type: cross Abstract: A proposed method for the control of groups of inspection spacecraft is Multi-Agent Reinforcement Learning (MA…
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots
arXiv:2607.02501v1 Announce Type: new Abstract: Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical de…
Cross-Platform Control for Autonomous Surface Vehicles via Adaptive Reinforcement Learning
arXiv:2607.02037v1 Announce Type: new Abstract: Autonomous surface vehicles vary widely in hydrodynamic and actuation characteristics, yet most controllers are …
A Convex Obstacle Avoidance Formulation
arXiv:2512.13836v2 Announce Type: replace-cross Abstract: Autonomous driving requires reliable collision avoidance in dynamic environments. Nonlinear Model Pred…
Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models
arXiv:2512.01715v2 Announce Type: replace Abstract: Vision-language-action (VLA) models that generate continuous action chunks via flow matching lack an interna…
PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
arXiv:2607.01938v1 Announce Type: new Abstract: Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodie…
Learning 3D-Gaussian Simulators from RGB Videos
arXiv:2503.24009v3 Announce Type: replace-cross Abstract: Realistic simulation is critical for applications ranging from robotics to animation. Learned simulato…
Robust Image Processing Techniques for Construction Environment Monitoring Using Underwater Robots
arXiv:2607.01915v1 Announce Type: cross Abstract: This paper proposes a robust image processing framework for underwater robot-based construction environment mo…
Adaptive Companionship for Group-Following Robots: Handling Dynamically Changing Group Formations
arXiv:2607.01287v1 Announce Type: new Abstract: Accompanying a group of humans is an essential aspect of developing human-like social cognition in robots. Howev…
Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration
arXiv:2603.06001v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural langua…
Regression Test Selection for Updated Capability Modules in Compositional ML Systems via Atomic-Quality Probes
arXiv:2604.26689v4 Announce Type: replace Abstract: Compositional machine-learning (ML) systems assemble runtime behavior from libraries of independently re-tra…
CommonRoad-Game: A Human-in-the-Loop Simulation Framework for Autonomous Driving
arXiv:2607.01382v1 Announce Type: new Abstract: Motion planning algorithms should be evaluated in human-in-the-loop environments to ensure they produce safe and…
NEUROSYMLAND: Neuro-Symbolic Landing-Site Assessment for Robust and Edge-Deployable UAV Autonomy
arXiv:2607.02277v1 Announce Type: new Abstract: Safe landing-site assessment in unstructured environments remains a key challenge for autonomous UAV deployment,…
VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer
arXiv:2512.11891v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in generalizing across diverse…
Imagining the Sense of Touch: Touch-Informed Manipulation via Imagined Tactile Representations
arXiv:2607.01684v1 Announce Type: new Abstract: Tactile sensing can substantially improve contact-rich robotic manipulation, yet its practical deployment remain…
Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving
arXiv:2604.03497v2 Announce Type: replace Abstract: Vision-language-model (VLM)-guided reinforcement learning (RL) has recently attracted significant attention …
Real-Time Visual Intelligence on Low-Cost UAVs: A Modular Approach for Tracking, Scanning, and Navigation
arXiv:2607.02298v1 Announce Type: new Abstract: Autonomous drones are rapidly transforming modern warfare and civil applications alike. This paper presents the …
CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
arXiv:2603.22435v2 Announce Type: replace Abstract: "Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) me…
Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs
arXiv:2607.02466v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- t…
LIME: Learning Intent-aware Camera Motion from Egocentric Video
arXiv:2607.02417v1 Announce Type: new Abstract: Autonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded …
Controllable Sim Agents with Behavior Latents
arXiv:2607.02496v1 Announce Type: new Abstract: Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpre…
Bridge-WA: Predicting Where and How the World Changes for Robotic Action
arXiv:2607.02195v1 Announce Type: new Abstract: General-purpose vision-language-action models benefit from large vision-language priors, but effective manipulat…
SPLC: Social Preference Learning for Crowd Robot Navigation
arXiv:2607.01925v1 Announce Type: new Abstract: Offline reinforcement learning (RL) holds significant potential for crowd robot navigation in human-robot coexis…
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
arXiv:2605.11020v2 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching…