Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesContact Modes Are Strata: What Geometric Structure Buys in Discrete-Continuous Planning
arXiv:2608.15541v1 Announce Type: new Abstract: Contact-rich manipulation poses a discrete question and a continuous one at once, namely which contacts are acti…
Pre-training Visual Dexterity in Simulation
arXiv:2608.15917v1 Announce Type: new Abstract: Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has la…
Learning Varying Physical Therapist-Patient Interactions for Robot-mediated Upper Limb Task-Specific Training
arXiv:2608.15995v1 Announce Type: new Abstract: Upper extremity motor function recovery is positively linked to Task-Specific Training (TST) and sufficient ther…
US-VLA: An Ultrasound Vision-Language-Action Model for Embodied Abdomina
arXiv:2608.16074v1 Announce Type: new Abstract: Artificial intelligence-assisted ultrasound scanning enhances diagnostic reliability and efficiency by providing…
Deep Probabilistic Indoor Gas Source Localization via Physical Dependency-Guided Sequential Inference
arXiv:2608.16221v1 Announce Type: new Abstract: Reliable gas source localization (GSL) is critical to safety in industrial and urban environments, yet remains c…
HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction
arXiv:2608.16222v1 Announce Type: new Abstract: Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically gro…
Robot-Body-Aware Traversal Risk Graph Planning for Wheeled-Legged Robots in Complex Terrain
arXiv:2608.16433v1 Announce Type: new Abstract: Traversal Risk Graphs (TRGs) provide a compact, terrain-aware representation for global navigation, but native T…
Exposing the Long-tail in Embodied Urban Navigation via Scalable Learning from In-the-Wild Videos
arXiv:2608.16476v1 Announce Type: new Abstract: Learning embodied urban navigation policies from real-world data is constrained by the cost of task-specific dat…
Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots
arXiv:2608.16686v1 Announce Type: cross Abstract: Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically…
Language-Guided Generation for Personalized Inspection Planning
arXiv:2506.02917v2 Announce Type: replace Abstract: We propose a training-free, Vision-Language Model (VLM)-guided approach for efficiently generating trajector…
Flow Motion Policy: Manipulator Motion Planning with Flow Matching Models
arXiv:2604.07084v2 Announce Type: replace Abstract: Open-loop end-to-end neural motion planners have recently been proposed to improve motion planning for robot…
Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis
arXiv:2510.08759v3 Announce Type: replace-cross Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is cruci…
Formalisms for Robotic Mission Specification and Execution: A Comparative Analysis
arXiv:2603.15427v2 Announce Type: replace-cross Abstract: Robots are increasingly deployed across diverse domains and designed for multi-purpose operation. As r…
Geometric Reconstruction of Extrinsic Contact Trajectories using Tactile Sensing and Proprioception for Tool Manipulation
arXiv:2606.22251v2 Announce Type: replace Abstract: Tactile sensing enables robots to perceive rich contact information at the grasp, supporting tasks such as o…
EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
arXiv:2605.21862v2 Announce Type: replace Abstract: Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on…
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
arXiv:2604.09860v4 Announce Type: replace Abstract: The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based bench…
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
arXiv:2512.23649v4 Announce Type: replace Abstract: Humans learn locomotion through visual observation, interpreting visual content first before imitating actio…
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
arXiv:2510.10903v2 Announce Type: replace Abstract: Embodied intelligence has witnessed remarkable progress in recent years, driven by advances in computer visi…
DiSA-IQL: Offline Reinforcement Learning for Robust Soft Robot Control under Distribution Shifts
arXiv:2510.00358v2 Announce Type: replace Abstract: Soft snake robots offer remarkable flexibility and adaptability in complex environments, yet their control r…
Relay-Based Coordination for Energy-Efficient Multi-Robot Pickup and Delivery
arXiv:2509.14127v3 Announce Type: replace Abstract: We consider the problem of delivering multiple packages from a single depot to distinct goal locations using…
X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization
arXiv:2608.16658v1 Announce Type: cross Abstract: Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding …
Trajectory-Level Automatic Curriculum Learning for Legged Locomotion on Unstructured Terrain
arXiv:2608.16164v1 Announce Type: cross Abstract: Training locomotion policies for complex unstructured terrain requires a curriculum to avoid early exploration…
Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
arXiv:2504.06961v2 Announce Type: replace Abstract: 3D assembly tasks, such as furniture assembly and component fitting, play a crucial role in daily life and r…
Adaptive Bridge: A Proxy-Based Decoupling Layer for Mitigating DDS Backpressure in ROS 2
arXiv:2608.15380v1 Announce Type: cross Abstract: In ROS 2 systems using DDS, a single slow subscriber on a RELIABLE topic can cause backpressure that degrades …
FollowUpBot: An LLM-Based Conversational Robot for Automatic Postoperative Follow-up
arXiv:2507.15502v1 Announce Type: cross Abstract: Postoperative follow-up plays a crucial role in monitoring recovery and identifying complications. However, tr…
Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation
arXiv:2608.16843v1 Announce Type: new Abstract: Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied a…
HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL
arXiv:2608.16837v1 Announce Type: new Abstract: Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist visi…
Semantic- and Density-Aware Planning for Accessibility-Preserving Multi-Object Placement
arXiv:2608.16741v1 Announce Type: new Abstract: Long-term manipulation planning requires robots to reason not only about immediate task success but also about h…
Design Optimization for Large High-Force Soft Robot Manipulators Under Gravitational Loads
arXiv:2608.16728v1 Announce Type: new Abstract: Designing large soft robots capable of generating high forces for physical human-robot interaction remains a sig…
Throwing a Tight Spiral American Football by a Humanoid Robot
arXiv:2608.16642v1 Announce Type: new Abstract: Accurate throwing of the American football requires precise regulation of release conditions, where coupled line…