Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesHoverflie: An empirical investigation of rotor shrouds to transform micro air vehicles into multi-modal hovercraft
arXiv:2608.06707v1 Announce Type: new Abstract: Small rotorcraft intended for use indoors or around the built environment have extremely limited flight duration…
R2S-EGO: Dual-Proxy Refinement for Sparse-Capture Real-to-Sim
arXiv:2608.06827v1 Announce Type: new Abstract: Real-to-sim (R2S) depends on scene representations that render observations along robot ego trajectories, yet de…
When Coordination Becomes a Threat: Communication Attacks in LLM-Controlled Multi-Robot Systems
arXiv:2608.06830v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used as high-level planners in embodied multi-robot systems, enabl…
Are Visual Place Recognition Models Recognizing Places or Conditions? Distractor-Augmented Evaluation and Condition Suppression
arXiv:2608.06847v1 Announce Type: new Abstract: Long-term Visual Place Recognition (VPR) is typically evaluated by matching queries from one condition against a…
Spatiotemporal Agility: Time-Constrained Reinforcement Learning for Vision-Guided Dynamic Quadrupedal Interception
arXiv:2608.06907v1 Announce Type: new Abstract: Legged robots require robust agility to perceive and interact with complex and dynamic environments within a con…
Cross-View Action Consistency for Camera-Robust Vision-Language-Action Policies
arXiv:2608.06965v1 Announce Type: new Abstract: Vision-language-action (VLA) policies fine-tuned from a fixed scene camera can fail when the camera is moved, ev…
Representation Handoffs for OpenArm-Based Laboratory Mobile Manipulation
arXiv:2608.07154v1 Announce Type: new Abstract: Open-source robotics and foundation models have lowered the barrier to embodied AI, yet language-guided laborato…
Depth-Wise Probing and Pruning of the Planning Token in a Driving Vision-Language-Action Model
arXiv:2608.07361v1 Announce Type: new Abstract: Vision-language-action (VLA) models route driving decisions through a deep language model, but it is unclear how…
Ising Acceleration for Multi-Robot Multi-Target Planning
arXiv:2608.06803v1 Announce Type: cross Abstract: Ising machines are emerging as promising hardware for combinatorial optimization. With recent advances in CMOS…
Vernata: Self-Supervised Learning of LiDAR Point Representations
arXiv:2608.06919v1 Announce Type: cross Abstract: LiDAR serves as a primary sensing modality for robots operating in outdoor environments. However, the performa…
Synthetic LiDAR Data Generation and Deterministic Downsampling for Point Cloud Classification on the Edge
arXiv:2608.07106v1 Announce Type: cross Abstract: Deploying three-dimensional deep learning frameworks to low-power embedded processors is bottlenecked by the u…
Learning to Walk With Less: A Dyna-Style Approach to Quadrupedal Locomotion
arXiv:2509.06296v2 Announce Type: replace Abstract: Traditional on-policy reinforcement learning (RL) controllers for quadrupedal locomotion often suffer from l…
Progress-Certified Reversible Simplex Supervision of Goal-Reaching Reinforcement Learning
arXiv:2601.19499v2 Announce Type: replace Abstract: Task completion is difficult to certify when state aggregation, model mismatch, and disturbances invalidate …
Accurate Trajectory Tracking with Model Predictive Contouring Control for Bird-Scale Flapping-Wing MAVs
arXiv:2605.06042v3 Announce Type: replace Abstract: Flapping-wing micro aerial vehicles offer quieter and safer operation than rotary-wing drones, yet achieving…
STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations
arXiv:2604.10055v3 Announce Type: replace Abstract: Despite their strong performance in embodied tasks, recent Vision-Language-Action (VLA) models remain highly…
EchoVLA: Robotic Vision-Language-Action Model with Synergistic Declarative Memory for Mobile Manipulation
arXiv:2511.18112v3 Announce Type: replace Abstract: Recent progress in Vision-Language-Action (VLA) models has enabled embodied agents to interpret multimodal i…
Rigid-Covert GNSS Spoofing of UAV Swarms: A Structural Blind Spot, Its Detection Limit, and Absolute-Anchor Defenses
arXiv:2608.06885v1 Announce Type: cross Abstract: Cooperative UAV-swarm defenses commonly cross-check GNSS positions against measured inter-drone geometry. We s…
Scalable Long-Horizon Planning with Staggered Updates for Lifelong MAPF
arXiv:2608.06702v1 Announce Type: cross Abstract: Lifelong Multi-Agent Path Finding (LMAPF) requires generating collision-free paths for large agent fleets unde…
Detection and Ranging of Transient Extrinsic Contacts Based on 6D Dynamic Tactile Sensing
arXiv:2608.07075v1 Announce Type: new Abstract: Delicate manipulation often involves transient and subtle collisions between a grasped object and the environmen…
AutoIntervene: Calibrated Intervention for Action-Chunking Imitation Learning Policies
arXiv:2608.07065v1 Announce Type: new Abstract: Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting sho…
C2Dex: Contact-Consistent Reconstruction and Retargeting for Dexterous Manipulation from Monocular Video
arXiv:2608.07045v1 Announce Type: new Abstract: High-quality demonstrations for dexterous robot manipulation are costly and difficult to collect, whereas monocu…
A Haptic Robot Finger Designed for Guqin Instrument Playing
arXiv:2608.07002v1 Announce Type: new Abstract: With the rapid advancement of humanoid robotics and embodied intelligence technologies, numerous musical instrum…
How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots
arXiv:2608.06898v1 Announce Type: new Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question…
Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection
arXiv:2608.06434v1 Announce Type: new Abstract: Embodied intelligence demands both long-horizon reasoning and real-time closed-loop responsiveness. Recent dual-…
Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models
arXiv:2608.06994v1 Announce Type: new Abstract: World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolutio…
M2-SMap: Memory-Efficient Semantic Mapping with Hierarchical Multi-Model Representation
arXiv:2608.07074v1 Announce Type: new Abstract: Dense point cloud maps, as a typically used mapping representation, are difficult to deploy on resource-constrai…
Panoramic Multimodal Semantic Occupancy Prediction for Quadruped Robots
arXiv:2603.13108v2 Announce Type: replace Abstract: Panoramic imagery provides holistic 360{\deg} visual coverage for environmental perception in quadruped robo…
Robot guide with multi-agent control and automatic scenario generation with LLM
arXiv:2509.10317v2 Announce Type: replace Abstract: The article describes the development of a hybrid social robot control architecture to overcome the limitati…
WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
arXiv:2608.07267v1 Announce Type: cross Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) in…
Beyond Visibility: Real-Time Surface Accessibility Fields from Sparse LiDAR
arXiv:2608.06412v1 Announce Type: cross Abstract: Understanding which surfaces in a scene are physically accessible to a given tool is fundamental for robotic i…