Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesSpherical-GOF: Geometry-Aware Panoramic Gaussian Opacity Fields for 3D Scene Reconstruction
arXiv:2603.08503v2 Announce Type: replace-cross Abstract: Omnidirectional images are increasingly used in robotics and vision due to their wide field of view. H…
Real-Time Model Checking for Closed-Loop Robot Reactive Planning
arXiv:2508.19186v2 Announce Type: replace Abstract: Reactive obstacle avoidance methods often cause agents to become trapped in local minima, because they can o…
Robust In-Hand Manipulation via Priors in Reinforcement Learning and Mechanical Design
arXiv:2607.12105v1 Announce Type: new Abstract: In-hand manipulation without external sensing is challenging due to uncertainties from finger-object contacts an…
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
arXiv:2507.15833v3 Announce Type: replace Abstract: Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions thr…
TADPO: Reinforcement Learning Goes Off-road
arXiv:2603.05995v2 Announce Type: replace Abstract: Off-road autonomous driving poses significant challenges such as navigating unmapped, variable terrain with …
Edge-Aware Thermal Infrared UAV Swarm Tracking
arXiv:2607.12544v1 Announce Type: cross Abstract: Thermal infrared (TIR) imaging is essential for UAV swarm operations in visually degraded environments. Howeve…
Flatness-Preserving Residual Learning for Real-Time Tight Quadrotor Formation Flight
arXiv:2607.12275v1 Announce Type: new Abstract: Quadrotors flying in tight formations are severely affected by turbulent aerodynamic interactions, such as downw…
Practical Judgment, Virtue, and Intuition in the Use of Opaque AI-Enabled Systems
arXiv:2607.12755v1 Announce Type: cross Abstract: AI-enabled systems are seeing increasing deployment across numerous domains, with many being "black boxes" wit…
Model-Based Diffusion Optimal Control for Multi-Robot Motion Planning
arXiv:2607.12423v1 Announce Type: new Abstract: Multi-Robot Motion Planning in continuous environments, where robots must generate dynamically feasible, collisi…
Contract-Grounded Behavior Tree Synthesis via Coding Agents
arXiv:2607.12220v1 Announce Type: new Abstract: Synthesizing deployable robot behavior trees (BTs) from natural language (NL) requires grounding to ensure every…
UR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress Proxies
arXiv:2607.12892v1 Announce Type: new Abstract: Modern robot learning systems increasingly rely on dense progress or value signals to evaluate intermediate stat…
EFLUX: Elastic Multi-Robot Formation Navigation and Adaptation with Agentic LLMs
arXiv:2607.12050v1 Announce Type: new Abstract: Multi-robot teams operating in confined or cluttered environments must adapt both their formation geometry and g…
Unveiling Complex Collective Behaviors from Simple Rewards
arXiv:2607.12861v1 Announce Type: new Abstract: Multi-agent Reinforcement Learning (MARL) holds great potential for robot swarms, but the black-box nature of ne…
More than a Manipulator: Planning Propellant-Free Attitude Maneuvers for Free-Floating Spacecraft
arXiv:2607.12130v1 Announce Type: new Abstract: Spacecraft attitude control is traditionally achieved using momentum exchange devices or propellant-consuming th…
Globalized Constrained Stein Variational Inference for Diverse Feasible Robot Motion Planning
arXiv:2607.12732v1 Announce Type: new Abstract: Robot motion planning is inherently multimodal, yet classical planners typically return only a single solution. …
VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation
arXiv:2607.12356v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a powerful end-to-end paradigm for robotic manipulation by m…
APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model
arXiv:2603.08862v2 Announce Type: replace Abstract: Autonomous navigation in highly constrained environments remains challenging for mobile robots. Classical na…
A Behavioral State Vocabulary in Sony ERS-111 R-CODE
arXiv:2607.12115v1 Announce Type: new Abstract: This paper presents a corpus-level analysis of generated behavior diagrams derived from Sony's R-CODE sample dis…
FlowWAM: Optical Flow as a Unified Action Representation for World Action Models
arXiv:2607.13017v1 Announce Type: new Abstract: World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action p…
MAMMOTH: A Multi-Modal End-to-End Policy for Off-Road Mobility Robust to Missing Modality
arXiv:2607.12965v1 Announce Type: new Abstract: Reliable autonomous navigation in unstructured off-road environments remains a critical unsolved challenge due t…
Infra-Swarm: Robust Vision-Based Multi-Robot Swarming via Near-Infrared Spectral Vision
arXiv:2607.12489v1 Announce Type: new Abstract: Distributed swarms typically rely on either active wireless communication or passive vision, and they are freque…
Diffusion Denoiser-Aided Gyrocompassing
arXiv:2507.21245v2 Announce Type: replace Abstract: An accurate initial heading angle is essential for efficient and safe navigation across diverse domains. Unl…
DiffRadar: Differentiable Physics-Aware Radar SLAM with Gaussian Fields
arXiv:2607.12265v1 Announce Type: new Abstract: Radar sensing is increasingly used in mobile systems because it operates reliably under poor lighting, adverse w…
RoboDesign1M: A Large-scale Dataset for Robot Design Understanding
arXiv:2503.06796v2 Announce Type: replace Abstract: Robot design is a complex and time-consuming process that requires specialized expertise. Gaining a deeper u…
Mind the Gap: Promises and Pitfalls of Hierarchical Planning in LeWorldModel
arXiv:2607.12547v1 Announce Type: new Abstract: We investigate whether temporal hierarchy can improve LeWorldModel on long-horizon goal-conditioned control. We …
DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter
arXiv:2602.05513v3 Announce Type: replace Abstract: Bimanual dexterous manipulation relies on integrating multimodal inputs to perform complex real-world tasks.…
ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning
arXiv:2607.12931v1 Announce Type: new Abstract: Reinforcement Learning (RL) has demonstrated significant potential for improving Vision-Language-Action (VLA) mo…
PixelLoop: Shortcut Topological Navigation with Pixel-Level Loops
arXiv:2607.12811v1 Announce Type: new Abstract: Although topological mapping and navigation have been studied extensively, the specific role and downstream effe…
LQG solution for POMDP without estimating states: A minimum variance approach
arXiv:2607.12135v1 Announce Type: cross Abstract: This paper investigates the control of discrete-time linear time-invariant (LTI) systems subject to incomplete…
LapSurgie: Humanoid Robots Performing Surgery via Teleoperated Handheld Laparoscopy
arXiv:2510.03529v3 Announce Type: replace Abstract: Robotic laparoscopic surgery has gained increasing attention in recent years for its potential to deliver mo…