Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesDemystifying When and Why VLAs Fail in Contact-Rich Tasks and How to Fix Them
arXiv:2608.01402v1 Announce Type: new Abstract: We address the problem of understanding when and why Vision-Language-Action models struggle with contact-rich ma…
Stress-Relief Annealing: Polynomial-Time Simulation-Free Layout Optimization for Automated Warehouses
arXiv:2608.01024v1 Announce Type: cross Abstract: We study the problem of optimizing physical layouts for automated warehouses, where hundreds to thousands of r…
RIT*: Riemannian Informed Trees for Cost-Adaptive Optimal Motion Planning
arXiv:2608.00822v1 Announce Type: new Abstract: We present Riemannian Informed Trees (RIT*), a planning framework that replaces Euclidean primitives in batch-in…
RL Bootstrapping of OpenVLA-OFT for a Novel Robot Embodiment
arXiv:2608.01013v1 Announce Type: new Abstract: Adapting a pretrained vision-language-action (VLA) policy to a new robot usually assumes embodiment-specific dem…
StochSIPP: Safe Interval Path Planning in Stochastic Dynamic Environments
arXiv:2608.00792v1 Announce Type: new Abstract: Safe navigation under uncertain time-dependent blockage requires anticipating observations before committing to …
From Digital to Physical Reservoir Computing: Co-Optimizing Soft Robotic Reservoirs via Dynamics Matching
arXiv:2608.00484v1 Announce Type: new Abstract: Soft robotic substrates are promising for Physical Reservoir Computing (PRC) because their compliant nonlinear d…
Any House Any Task: Scalable Long-Horizon Planning for Abstract Human Tasks
arXiv:2602.12244v2 Announce Type: replace Abstract: Open world language conditioned task planning is crucial for robots operating in large-scale household envir…
SynAgent: Generalizable Cooperative Humanoid Manipulation via Solo-to-Cooperative Agent Synergy
arXiv:2604.18557v2 Announce Type: replace-cross Abstract: Controllable cooperative humanoid manipulation is a fundamental yet challenging problem for embodied i…
RF-HOI: Recognize Human-Object Interaction with Radio Frequency Signals
arXiv:2608.00289v1 Announce Type: cross Abstract: Recognizing Human-Object Interactions (HOI) is essential for intelligent systems, underpinning applications in…
Learning to Predict Contact Force Distributions from Vision Leveraging Object Geometry Priors
arXiv:2608.00464v1 Announce Type: new Abstract: Based on vision and prior experience, humans can make rough physical predictions and adjust their manipulation s…
LOCUS-DT: Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins
arXiv:2608.00406v1 Announce Type: cross Abstract: Accurate indoor localization is essential for emerging applications in robotic navigation and search and rescu…
Perception-and-action system for humanoid robot task execution in construction
arXiv:2608.01600v1 Announce Type: new Abstract: Humanoid robots, with their human-like shape and multi-tasking capabilities, are well-aligned with human-dominat…
Uncovering and Mitigating Positional Blind Spots in Vision-Language-Action Models
arXiv:2608.01573v1 Announce Type: new Abstract: Recent Vision-Language-Action (VLA) models achieve promising performance in robotic manipulation, typically meas…
Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions
arXiv:2604.15395v2 Announce Type: replace Abstract: Over the recent years, the field of robotics has been undergoing a transformative paradigm shift from fixed,…
TS-MAMP: A Remanufactured Agricultural Robot Powered by Second-Life EV Components and NMS-Free On-Device Weed Detection
arXiv:2608.02270v1 Announce Type: new Abstract: Agriculture 4.0 robotic systems improve field efficiency yet remain too capital-intensive for the fragmented sma…
Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State Models
arXiv:2409.07163v3 Announce Type: replace Abstract: Diffusion models have been widely employed in the field of 3D manipulation due to their efficient capability…
Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models
arXiv:2608.02497v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models excel in robotic manipulation but suffer catastrophic performance drops when…
TANGO-VIO: Triangulation-Aware Navigation with Guaranteed Feature-Observability for Visual-Inertial Odometry
arXiv:2608.02079v1 Announce Type: new Abstract: In vision-aided navigation and visual-inertial odometry, the quality of triangulated three-dimensional feature p…
Complete Motion Planning using Workspace-Fibered Decomposition for nR-Planar Manipulator
arXiv:2608.01172v1 Announce Type: new Abstract: We propose a workspace-fibered decomposition framework for motion planning in nR planar redundant manipulators o…
Learning Smooth SE(3) Trajectories under Left-Invariant Riemannian Metrics
arXiv:2608.01562v1 Announce Type: new Abstract: Optimal trajectory generation for rigid-body motions on Lie groups can be formulated as a variational problem th…
STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision
arXiv:2608.01535v1 Announce Type: cross Abstract: Vision-language models (VLMs) are emerging as a key component of embodied intelligence, with growing applicati…
A Robotic System for Automated Manufacturing of Dielectric Elastomer Actuators
arXiv:2608.00369v1 Announce Type: new Abstract: This letter presents an automated robotic manufacturing system for soft capacitors which operate as actuators an…
Residual-Based Adaptive Kalman Filtering for Legged Robot State Estimation
arXiv:2608.02316v1 Announce Type: new Abstract: State estimation is a key component in model-based control of walking robots and, more broadly, applicable where…
Localization in Spatiotemporal Fields via Environmental PDEs
arXiv:2608.00272v1 Announce Type: new Abstract: This paper proposes a localization framework that uses spatiotemporal fields governed by partial differential eq…
From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving
arXiv:2602.10719v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) driving augments end-to-end (E2E) planning with language-enabled visual backbon…
Motion Planning for Mobile Manipulators Navigating Doorways via Model Predictive Control
arXiv:2608.00206v1 Announce Type: new Abstract: Navigating doorways is a fundamental capability for mobile manipulators operating in human environments, requiri…
FeDepth: Federated Learning for Depth Estimation under Robot Heterogeneity
arXiv:2608.01129v1 Announce Type: new Abstract: Although recent robot perception research emphasizes training on data from diverse environments to improve gener…
GenTrack: Physical Alignment for Robot-Native Motion Generation and Zero-Shot Humanoid Tracking
arXiv:2608.01410v1 Announce Type: new Abstract: General-purpose humanoid trackers can execute diverse references, but their zero-shot coverage depends on large …
MemoAct: Atkinson-Shiffrin-Inspired Hierarchical Memory-Augmented Policy for Robotic Manipulation
arXiv:2603.18494v2 Announce Type: replace Abstract: Memory-augmented robotic policies are essential in handling memory-dependent tasks. However, existing approa…
Long-Term Memory for VLA-based Agents in Open-World Task Execution
arXiv:2604.15671v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated significant potential for embodied decision-making; ho…