Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesComparing Commercial Depth Sensor Accuracy for Medical Applications
arXiv:2606.13028v1 Announce Type: new Abstract: Depth estimation has numerous medical and surgical applications. We benchmark four depth sensors on a porcine bo…
SERF: Spatiotemporal Environment and Robot Feature Map for Long-Horizon Mobile Manipulation
arXiv:2606.12956v1 Announce Type: new Abstract: Long-horizon robot mobile manipulation requires continual reasoning about localization, environment changes, and…
Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning
arXiv:2606.12910v1 Announce Type: new Abstract: For robotics to be effectively integrated into household or industrial environments, machines must adapt to natu…
Learning to Adapt: Representation-Based Reinforcement Learning for Multi-Task Skill Transfer
arXiv:2606.12890v1 Announce Type: new Abstract: Reinforcement learning has achieved remarkable success in learning complex control policies, yet its applicabili…
AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots
arXiv:2606.12859v1 Announce Type: new Abstract: Aerial manipulation systems have long suffered from representation coupling in end-to-end control, as platform-l…
EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations
arXiv:2606.12604v1 Announce Type: new Abstract: Dexterous manipulation is limited by the cost of collecting large-scale robot demonstrations. Egocentric human v…
Active Semantic Perception
arXiv:2510.05430v2 Announce Type: replace Abstract: We develop an approach for active semantic perception, which refers to using the semantics of the scene for …
ReactEMG Stroke: Healthy-to-Stroke Few-shot Adaptation for sEMG-Based Intent Detection
arXiv:2601.22090v2 Announce Type: replace Abstract: Surface electromyography (sEMG) is a promising control signal for assist-as-needed hand rehabilitation after…
SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models
arXiv:2602.04208v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robotic control…
Adaptive-Horizon Conflict-Based Search for Closed-Loop Multi-Agent Path Finding
arXiv:2602.12024v2 Announce Type: replace Abstract: MAPF is a core coordination problem for large robot fleets in automated warehouses and logistics. Existing a…
AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly
arXiv:2604.08983v2 Announce Type: replace Abstract: Spatial reasoning is a fundamental capability for embodied intelligence, especially for fine-grained manipul…
From Digital to Physical: Digital Agents as Autonomous Coaches for Physical Intelligence
arXiv:2601.21570v2 Announce Type: replace-cross Abstract: The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fuele…
Triangle Splatting SLAM
arXiv:2605.31419v2 Announce Type: replace-cross Abstract: We present a dense RGB-D SLAM system using differentiable triangles as the 3D map representation. Whil…
Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
arXiv:2606.01621v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have become a common foundation for vision-and-language navigation in co…
From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
arXiv:2507.22028v2 Announce Type: replace-cross Abstract: Navigation foundation models trained on massive web-scale data enable agents to generalize across dive…
Action-Effect Memory Pretraining for Robot Manipulation
arXiv:2606.12499v1 Announce Type: new Abstract: We present AEM, an Action-Effect Memory pretraining framework for robot manipulation that learns compact tempora…
EgoMoD: Predicting Global Maps of Dynamics from Local Egocentric Observations
arXiv:2603.00167v2 Announce Type: replace Abstract: Efficient navigation in dynamic environments requires anticipating how motion patterns evolve beyond the rob…
Miniature Testbed for Validating Multi-Agent Cooperative Autonomous Driving
arXiv:2511.11022v2 Announce Type: replace Abstract: Cooperative autonomous driving, which extends vehicle autonomy by enabling real-time collaboration between v…
GLIDE: A Coordinated Aerial-Ground Framework for Search and Rescue in Unknown Environments
arXiv:2509.14210v4 Announce Type: replace Abstract: We present a cooperative aerial-ground search-and-rescue (SAR) framework that pairs two unmanned aerial vehi…
Data-Driven Soft Robot Control via Adiabatic Spectral Submanifolds
arXiv:2503.10919v3 Announce Type: replace Abstract: The mechanical complexity of soft robots creates significant challenges for their model-based control. Speci…
PolyFlow: Safe and Efficient Polytope-Constrained Flow Matching with Constraint Embedding and Projection-free Update
arXiv:2606.13400v1 Announce Type: cross Abstract: While flow-based generative models have demonstrated strong performance across a wide range of domains, deploy…
Visual Place Recognition in Forests with Depth-Aware Distillation
arXiv:2606.13206v1 Announce Type: cross Abstract: Visual place recognition in natural forest environments remains challenging due to repetitive vegetation, weak…
MPC for underactuated spacecraft control with a Lyapunov supervised physics-informed neural network correction layer
arXiv:2606.13113v1 Announce Type: cross Abstract: Underactuated spacecraft faces controllability limitations and heightened sensitivity to environmental disturb…
Effects of Social Interactions in Self-Organising Railway Traffic Management
arXiv:2606.13068v1 Announce Type: cross Abstract: Recent research is exploring self-organised traffic management as a solution for scaling to complex real-world…
TrajGenAgent: A Hierarchical LLM Agent for Human Mobility Trajectory Generation
arXiv:2606.12657v1 Announce Type: cross Abstract: Human mobility data is important for transportation, urban planning, and epidemic control, but large-scale tra…
Individual Control Barrier Functions-Guided Diffusion Model for Safe Offline Multi-Agent Reinforcement Learning
arXiv:2606.12640v1 Announce Type: cross Abstract: Offline reinforcement learning allows control policies to be learned directly from data without online interac…
WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning
arXiv:2604.08958v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) in robotics is often limited by the cost and risk of data collection, moti…
GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
arXiv:2510.03896v2 Announce Type: replace-cross Abstract: Vision-language models demonstrate strong reasoning and planning abilities, yet grounding these predic…
DiffCoord: Differentiable Coordination for Distributed Multi-Agent Trajectory Optimization
arXiv:2509.01630v3 Announce Type: replace-cross Abstract: Integrating the Alternating Direction Method of Multipliers (ADMM) with Differential Dynamic Programmi…
Safety Case Patterns for VLA-based driving systems: Insights from SimLingo
arXiv:2603.16013v3 Announce Type: replace Abstract: Vision-Language-Action (VLA)-based driving systems represent a significant paradigm shift in autonomous driv…