Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesFrom Ad Hoc Pilots to Repeatable Patterns: Structuring Drone Collaboration in Emergency Services with DroneLets
arXiv:2606.17839v1 Announce Type: new Abstract: Drones hold promise for supporting emergency services, but their integration into workflows remains ad hoc and c…
PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space
arXiv:2606.17924v1 Announce Type: new Abstract: Current Vision-Language-Action (VLA) models face a trade-off between efficient action generation and explicit de…
Uncertainty Quantification for Flow-Based Vision-Language-Action Models
arXiv:2606.18043v1 Announce Type: new Abstract: Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads t…
A Hybrid Optimization Framework for Grasp Synthesis under Partial Observations
arXiv:2606.18053v1 Announce Type: new Abstract: We propose a hybrid grasp synthesis framework that combines a learning-based Energy-Based Model (EBM) with an an…
Beyond Failure Recovery: An Engagement-Aware Human-in-the-loop Framework for Robotic Systems
arXiv:2606.18189v1 Announce Type: new Abstract: Conventional human-in-the-loop approaches typically involve users only when a robot encounters failure or uncert…
EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies
arXiv:2606.18239v1 Announce Type: new Abstract: We present EBench, a simulation benchmark that diagnoses generalist mobile manipulation policies beyond a single…
Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
arXiv:2606.18247v1 Announce Type: new Abstract: Robots deployed in the real world should learn from their experience and improve over time. This requires a mech…
Credibility-Weighted Pricing of Autonomous Vehicle Liability Under Operational Design Domain Shift
arXiv:2606.17451v1 Announce Type: cross Abstract: Automated Driving System deployments create a foundational ratemaking challenge: sparse experience, shifting o…
Adaptive Volumetric Mechanical Property Fields Invariant to Resolution
arXiv:2606.18231v1 Announce Type: cross Abstract: Accurate mechanical properties (or materials) Young's modulus ($E$), Poisson's ratio ($\nu$) and density ($\rh…
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
arXiv:2605.31286v2 Announce Type: replace Abstract: Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable…
SCC-Loc: A Unified Semantic Cascade Consensus Framework for UAV Thermal Geo-Localization
arXiv:2604.03120v2 Announce Type: replace-cross Abstract: Cross-modal Thermal Geo-localization (TG) provides a robust, all-weather solution for Unmanned Aerial …
ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control
arXiv:2606.03177v2 Announce Type: replace Abstract: Human demonstrations provide strong priors for robot manipulation, yet it is non-trivial to transfer them to…
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
arXiv:2603.22281v2 Announce Type: replace-cross Abstract: Recent progress in latent world models (e.g., V-JEPA2) has shown promising capability in forecasting f…
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
arXiv:2603.03485v3 Announce Type: replace-cross Abstract: Recent video diffusion models have achieved impressive capabilities as large-scale generative world mo…
Critique of World Model: A Generative Latent Prediction Architecture for World Modeling
arXiv:2507.05169v4 Announce Type: replace-cross Abstract: World Model, the algorithmic simulator of the real-world environment which biological agents experienc…
Any2Any: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking
arXiv:2605.23733v2 Announce Type: replace Abstract: Whole-body tracking (WBT) models have become a key foundation for humanoid robots, enabling them to imitate …
When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning
arXiv:2605.05172v2 Announce Type: replace Abstract: Behavior Cloning (BC) has emerged as a highly effective paradigm for robot learning. However, BC lacks a sel…
LAGO Policy: Latency-Aware Asynchronous Diffusion Policies with Goal-Directed Collision-Free Planning for Smooth Manipulation
arXiv:2606.17982v1 Announce Type: new Abstract: Diffusion-based visuomotor policies deployed with asynchronous inference often exhibit inter-chunk discontinuiti…
Physical Imitation Learning: Distilling Control Policies into Passive Elasticity
arXiv:2604.00611v2 Announce Type: replace Abstract: Due to brain-body co-evolution, animals' intrinsic body dynamics play a crucial role in their energy-efficie…
RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models
arXiv:2506.17639v2 Announce Type: replace Abstract: Vision-Language-Action models (VLA) have demonstrated remarkable capabilities and strong potential in comple…
K-VARK: Kernelized Variance-Aware Residual Kalman Filter for Sensorless Force Estimation in Collaborative Robots
arXiv:2512.13009v2 Announce Type: replace Abstract: Reliable estimation of contact forces is crucial for ensuring safe and precise interaction of robots with un…
ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation
arXiv:2606.17937v1 Announce Type: new Abstract: Most Vision-Language-Action (VLA) models map observations directly to actions without explicit reasoning, limiti…
WAM-RL: World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFT
arXiv:2606.17906v1 Announce Type: new Abstract: Recent World-Action (WA) models demonstrate strong generalization ability and data efficiency, but they typicall…
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
arXiv:2606.17200v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajec…
Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
arXiv:2606.17846v1 Announce Type: new Abstract: Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data und…
FLAP: FOV-Constrained Active Perception Planning for Prior-Map-Free 3D Navigation
arXiv:2606.17630v1 Announce Type: new Abstract: Safe and efficient trajectory planning in unknown, cluttered 3D environments constitutes a critical bottleneck f…
GASE: Gaussian Splatting-Based Automated System for Reconstructing Embodied-Simulation Environments
arXiv:2606.17520v1 Announce Type: new Abstract: Training embodied agents in the real world requires skilled operators and expensive hardware. Simulation environ…
ParkingTransformer: LLM-Enhanced End-to-End Trajectory Planning for Autonomous Parking
arXiv:2606.17082v1 Announce Type: new Abstract: End-to-end autonomous parking has emerged as a critical task within the realm of autonomous driving. However, ex…
MagicSim: A Unified Infrastructure for Executable Embodied Interaction
arXiv:2606.17511v1 Announce Type: new Abstract: Robot learning and embodied agents now require simulation to serve as a shared execution substrate linking contr…
MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation
arXiv:2606.17598v1 Announce Type: new Abstract: Humans naturally leverage diverse sensing modalities to interact with the physical world, while most Vision-Lang…