Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesAct, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models
arXiv:2603.05147v2 Announce Type: replace-cross Abstract: Current research on Vision-Language-Action (VLA) models predominantly focuses on enhancing generalizat…
FloAff-Kitchen: Bridging Navigation and Manipulation via Canonical and Progressive Floor Affordance Learning
arXiv:2607.24207v1 Announce Type: new Abstract: Mobile manipulation requires robots to identify Floor Affordance (FloAff) that maximizes downstream manipulation…
Try Once, Then Optimal: De-Redundified Procedure Memory for Cross-Episode Exploration Amortization
arXiv:2607.23702v1 Announce Type: new Abstract: Manipulating objects with hidden internal state, such as a latched microwave, forces a robot to probe before it …
Co-planning of Flight Corridors and Communication Infrastructure for Urban Drone Logistics Networks
arXiv:2607.23989v1 Announce Type: new Abstract: Reliable wireless connectivity is essential for urban air mobility (UAM) networks in dense urban environments. I…
Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance
arXiv:2607.22667v1 Announce Type: cross Abstract: This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for single-ve…
Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAG
arXiv:2607.24313v1 Announce Type: cross Abstract: Marine life monitoring is limited by strict energy constraints, poor underwater connectivity, and the high cos…
Learning-based Hierarchical Tracheal Anatomy Understanding from Sparse Surgical Demonstration Annotations for Ultrasound Robots
arXiv:2607.22789v1 Announce Type: cross Abstract: Tracheostomy requires precise localization of the tracheal incision site; however, conventional manual palpati…
CReF: Cross-modal and Recurrent Fusion for Depth-conditioned Humanoid Locomotion
arXiv:2603.29452v3 Announce Type: replace Abstract: Stable traversal over geometrically complex terrain increasingly requires exteroceptive perception, yet prio…
Pose-Aware Modeling to Mitigate Pose-Related Artifacts in Tactile Gloves
arXiv:2607.22964v1 Announce Type: new Abstract: Tactile gloves digitize contact and force during hand-object interactions, enabling robotics applications in dex…
Stress-testing large language model agents in a robotic chemistry laboratory
arXiv:2607.23045v1 Announce Type: cross Abstract: AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable phys…
Structured Observation Language for Efficient and Generalizable Vision-Language Navigation
arXiv:2603.27577v2 Announce Type: replace-cross Abstract: Vision-Language Navigation (VLN) requires an embodied agent to navigate complex environments by follow…
AutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models
arXiv:2603.28963v2 Announce Type: replace Abstract: Simulation with realistic traffic agents is essential for validating autonomous driving systems. Existing da…
DeReCo: Decoupling Representation and Coordination Learning for Object-Adaptive Decentralized Multi-Robot Cooperative Transport
arXiv:2603.08111v2 Announce Type: replace Abstract: Generalizing decentralized multi-robot cooperative transport across objects with diverse shapes and physical…
Error-State LQR Formulation for Quadrotor UAV Trajectory Tracking
arXiv:2501.15768v3 Announce Type: replace Abstract: This article presents an error-state Linear Quadratic Regulator (LQR) formulation for robust trajectory trac…
Can Context Bridge the Reality Gap? Sim-to-Real Transfer of Context-Aware Policies
arXiv:2511.04249v2 Announce Type: replace Abstract: Sim-to-real transfer remains a major challenge in reinforcement learning (RL) for robotics, as policies trai…
SLAM-Former: Putting SLAM into One Transformer
arXiv:2509.16909v2 Announce Type: replace-cross Abstract: We present SLAM-Former, a neural approach that integrates full SLAM capabilities into a single transfo…
Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction
arXiv:2602.00575v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) provides a promising pathway for continuously advancin…
LEACL: LLM-Enhanced Automatic Curriculum Learning for Reinforcement Learning in Long-Horizon Manipulation Tasks
arXiv:2607.23515v1 Announce Type: new Abstract: Long-horizon manipulation tasks pose significant challenges for reinforcement learning due to sparse reward sign…
HELIOS: An LLM-Driven Autonomous Indirect Trajectory Optimization Agent
arXiv:2607.24051v1 Announce Type: cross Abstract: Low-thrust trajectory optimization is a core technology in deep-space mission design. Indirect methods based o…
$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
arXiv:2607.23783v1 Announce Type: new Abstract: We present $N_0$-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both futu…
SILICA: Repurposing Diffusion Priors for Joint Glass Segmentation and Depth Estimation
arXiv:2607.24249v1 Announce Type: cross Abstract: Standard depth sensors systematically fail on transparent surfaces, creating corrupted 3D maps and severe navi…
A Few Words Go a Long Way: Language Guided Robot Policy Synthesis
arXiv:2607.23784v1 Announce Type: new Abstract: While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remai…
UNet: A Generic and Reliable Multi-UAV Communication and Networking Architecture for Heterogeneous Applications
arXiv:2411.03048v2 Announce Type: replace-cross Abstract: The rapid growth of UAV applications necessitates a robust communication and networking system archite…
A Cyclic Adaptation-Generalization Framework with Uncertainty-Guided Self-Paced Learning for Long-Term Brain-Machine Interfaces
arXiv:2607.24031v1 Announce Type: cross Abstract: Brain-Machine Interfaces (BMIs), which link the brain to external devices, hold great potential in rehabilitat…
Moving-Horizon Estimation and Nonlinear Model Predictive Control of Cable-Driven Soft Manipulators
arXiv:2607.24029v1 Announce Type: new Abstract: Precise control of soft manipulators remains challenging due to the difficulty of developing accurate yet comput…
BC-NMPC: Battery-Constrained NMPC with Propulsion Prediction and Replanning for High-Speed Flight
arXiv:2607.23867v1 Announce Type: new Abstract: Trajectory tracking performance of Uncrewed Aerial Vehicles (UAVs) degrades during high-speed and agile flight d…
WCM: World-Cognition Model for Generalizable Human-Robot Interaction
arXiv:2607.22999v1 Announce Type: new Abstract: Language agents can now interact fluently with users in software, but robots still struggle to bring comparable …
Online Lidar-Only Odometry with Retrospective Refinement of Overlapping Submaps
arXiv:2503.21293v3 Announce Type: replace Abstract: Lidar odometry aims to estimate the ego-motion of a mobile platform from a stream of lidar scans. Traditiona…
PAC-DP: PAC-Bayesian Diffusion Policy Learning
arXiv:2607.24296v1 Announce Type: new Abstract: Diffusion Policies (DPs) are able to perform complex manipulation tasks. However, DPs are typically trained by m…
Towards Ultrafast Depth Sensing Via Active Event-based Stereo Vision
arXiv:2607.23684v1 Announce Type: new Abstract: Conventional frame-based imaging for active stereo systems has encountered major challenges in fast-motion scena…