Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesOctopus Protocol: One-Shot Hardware Discovery and Control for AI Agents via Infrastructure-as-Prompts
arXiv:2605.09055v2 Announce Type: replace Abstract: Bringing a previously unintegrated device under the control of an AI agent still requires device-specific en…
YOLO with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection with trustworthy multimodal AI in computer vision perception
arXiv:2603.23037v2 Announce Type: replace-cross Abstract: The trustworthy object detection capabilities of a novel Kolmogorov-Arnold network framework are exami…
LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory
arXiv:2609.02350v2 Announce Type: replace-cross Abstract: Vision-Language Navigation (VLN) requires an embodied agent to follow natural-language instructions in…
SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control
arXiv:2605.22894v3 Announce Type: replace-cross Abstract: Controlling physics-based humanoids from natural-language instructions is a critical step toward gener…
Out-of-Distribution Semantic Occupancy Prediction
arXiv:2506.21185v3 Announce Type: replace-cross Abstract: 3D semantic occupancy prediction is crucial for autonomous driving, providing a dense, semantically ri…
Multi-Robot Bearing-based Pose Estimation via Angle Rigidity
arXiv:2606.03931v2 Announce Type: replace Abstract: This letter proposes a novel distributed pose estimator for multi-robot systems evolving on $\mathrm{SE}(3)$…
Risk-Aware Optimal Control with Rulebooks
arXiv:2609.05199v1 Announce Type: cross Abstract: We consider safety-critical control problems involving multiple requirements with different priorities and unc…
Pack It My Way: Triadic Human-Robot Collaboration for Personalized Autonomous Packing
arXiv:2609.04620v1 Announce Type: new Abstract: Personalized autonomous packing requires robots to account for resident preferences that cannot be inferred from…
Open-Set 3D Scene Graphs for Field Robotics: An Outdoor Case Study
arXiv:2609.04607v1 Announce Type: new Abstract: Three-dimensional scene graphs (3DSGs) have emerged as a promising approach for building geometrically grounded,…
Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI
arXiv:2609.04552v1 Announce Type: new Abstract: Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with huma…
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
arXiv:2609.04355v1 Announce Type: new Abstract: Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demandin…
Scalable Edge-assisted Fusion and Path Prediction for Connected Autonomous Vehicles
arXiv:2609.04364v1 Announce Type: new Abstract: The planning algorithms inside an Autonomous Vehicle (AV) rely on information from on-board sensors whose line o…
Dressing in Motion: A Human Motion-Aware Diffusion Policy for Robot-Assisted Dressing
arXiv:2609.04759v1 Announce Type: new Abstract: Robotic dressing assistance is a promising solution for supporting older adults with physical impairments in dai…
Coupled Control and Wireless World Models for Resilient Remote Robotic Control
arXiv:2609.04851v1 Announce Type: new Abstract: Remote robotic systems operating over wireless networks must maintain reliable control despite limited communica…
One Word, Different Action: A Real-Robot Benchmark for Language-Conditioned Embodied Reasoning
arXiv:2609.05260v1 Announce Type: new Abstract: Natural-language instruction changes can directly alter robot behavior. A reliable embodied system should preser…
Temporal Tactile Encoding and Compliance for Intent-Aware Robot-to-Human Bimanual Handover
arXiv:2609.05282v1 Announce Type: new Abstract: Reliable robot-to-human handover requires the robot to infer when the person is ready to receive the object, and…
Human-Human & Human-Robot Interaction Transformer (H2INT) for Robot Navigation in Dense and Uncertain Crowds
arXiv:2609.05300v1 Announce Type: new Abstract: Safe robot navigation in dense crowds requires reasoning about pedestrian motion and how it may change in respon…
What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies
arXiv:2609.05376v1 Announce Type: new Abstract: Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when…
Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models
arXiv:2609.05401v1 Announce Type: new Abstract: Vision-language models are increasingly used as reward functions for robotic learning, but this role requires pa…
Game-Theoretic Drone Swarm Defense: A Case Study in Applied Differential Game Theory
arXiv:2609.04394v1 Announce Type: cross Abstract: This technical report is a study of the use of differential game (DG) theory to solve the target-assignment an…
RedVLA: Physical Red Teaming for Vision-Language-Action Models
arXiv:2604.22591v2 Announce Type: replace Abstract: The real-world deployment of Vision-Language-Action (VLA) models remains limited by the risk of unpredictabl…
X2-N: A Transformable Wheel-legged Humanoid Robot with Dual-mode Locomotion and Manipulation
arXiv:2604.21541v2 Announce Type: replace Abstract: Wheel-legged robots combine the efficiency of wheeled locomotion with the versatility of legged systems, ena…
Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
arXiv:2603.25685v2 Announce Type: replace Abstract: Action-conditioned robot world models generate future video frames of the manipulated scene given a robot ac…
CoFreeVLA: Short-Horizon Collision-Free Dual-Arm Manipulation via Vision-Language-Action Model and Risk Estimation
arXiv:2601.21712v3 Announce Type: replace Abstract: Vision Language Action (VLA) models enable instruction-following manipulation, yet their deployment on coord…
A Biomimetic Vertebraic Soft Robotic Tail for High-Speed, High-Force Dynamic Maneuvering
arXiv:2509.20219v2 Announce Type: replace Abstract: Robotic tails can enhance the stability and maneuverability of mobile robots, but current designs face a tra…
Squint: Fast Visual Reinforcement Learning for Sim-to-Real Robotics
arXiv:2602.21203v2 Announce Type: replace Abstract: Visual reinforcement learning is appealing for robotics but expensive. Off-policy methods are sample-efficie…
Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation
arXiv:2609.05369v1 Announce Type: new Abstract: Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon pr…
Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction
arXiv:2609.05361v1 Announce Type: new Abstract: Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in real-…
Adaptation Needs in Robotic Systems: Assessing Behavior Trees and Their Enhancement
arXiv:2609.05331v1 Announce Type: new Abstract: Robotic systems increasingly operate in dynamic, uncertain, and open-ended environments, where design-time assum…
FIRE-LIVWO: Robust LiDAR-Inertial-Visual-Wheel Odometry via Failure-Immune mmWave Radar Enhancement
arXiv:2609.05325v1 Announce Type: new Abstract: Achieving robust SLAM in large-scale underground coal mines with complex structures and severe degeneracies rema…