Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesProtoAct: Turning Wet-Lab Protocols into Embodied Robotic Actions
arXiv:2608.01690v1 Announce Type: new Abstract: Biological wet-lab protocols are written for trained researchers and often leave routine operations, state-depen…
Hermite Curves as Trajectory Priors for Vision-Language-Action Models
arXiv:2608.01265v1 Announce Type: new Abstract: Despite recent progress in Vision-Language-Action (VLA) models for robotic manipulation, the action chunk remain…
ReTouch: Empowering Contact-Rich Dexterous Manipulation with Online-Refined Tactile Prediction
arXiv:2608.01824v1 Announce Type: new Abstract: Fusing tactile signals has proven effective for contact-rich manipulation, enabling robots to perceive contact s…
MROPE: A Multi-Robot Safe Cooperative Strategy via combined Predictive Safety Filters and Ellipse-based Constraint Compression
arXiv:2607.29203v1 Announce Type: new Abstract: Deploying drone swarms to track a dynamic target in cluttered environments presents severe computational and saf…
D-VLC: Decentralized Vision-Language Collaboration for Heterogeneous Embodied Multi-Robot Systems in Unknown Environments
arXiv:2607.29009v1 Announce Type: new Abstract: Multi-robot systems, particularly heterogeneous robot swarms, can improve the efficiency of complex task executi…
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
arXiv:2607.29613v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for ro…
RayViT: Ray-Conditioned Visual Representations for Viewpoint-Robust Imitation Learning
arXiv:2607.29622v1 Announce Type: new Abstract: Visual imitation learning enables robots to acquire visuomotor skills directly from images, yet RGB observations…
Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events
arXiv:2509.25146v2 Announce Type: replace-cross Abstract: This paper develops a mathematical argument and algorithms for building representations of data from e…
Kinodynamic Motion Retargeting for Humanoid Locomotion via Multi-Contact Whole-Body Trajectory Optimization
arXiv:2603.09956v2 Announce Type: replace Abstract: We present the KinoDynamic Motion Retargeting (KDMR) framework, a novel approach for humanoid locomotion tha…
CorrelationFlow: A Training-Free Geometric Approach for LiDAR Scene Flow Estimation
arXiv:2607.29237v1 Announce Type: cross Abstract: LiDAR scene flow estimation has settled into a monoculture: nearly all recent methods share the same feed-forw…
AquaJEPA: Action-Conditioned Multimodal Predictive Representations for Underwater Robot Dynamics
arXiv:2607.29393v1 Announce Type: new Abstract: Underwater robots combine complementary sensors whose reliability changes abruptly with water visibility, viewpo…
Tri-Space Operational Control of Redundant Multilink and Hybrid Cable-Driven Parallel Robots Using an Iterative-Learning based Reactive Approach
arXiv:2607.29500v1 Announce Type: new Abstract: Cable-Driven Parallel Robots (CDPRs) are a type of parallel mechanism in which cables are used as actuators. Due…
BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning
arXiv:2607.29302v1 Announce Type: new Abstract: Reliable robot learning requires a world simulator that can predict action consequences before execution on phys…
Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints
arXiv:2501.04426v2 Announce Type: replace-cross Abstract: Offline diversity maximization under imitation constraints can transform demonstration data into a set…
Combining Large Language Models and Symbolic Reasoning for Multi-Robot Temporal Planning through Explainable Knowledge Bases
arXiv:2502.19135v2 Announce Type: replace-cross Abstract: We present PLANTOR, a framework for generating and executing multi-robot task plans from natural-langu…
Vision-based Goal-Reaching Control for Mobile Robots Using a Hierarchical Learning Framework
arXiv:2601.00610v2 Announce Type: replace Abstract: Reinforcement learning (RL) has strong potential in robotics, but exploration-based training complicates saf…
GPA-RAM: Grasp-Pretraining Augmented Robotic Attention Mamba for Spatial Task Learning
arXiv:2504.19683v4 Announce Type: replace Abstract: Fine-grained robotic manipulation often fails when inaccurate initial grasps propagate errors and necessitat…
FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Sampling
arXiv:2607.29596v1 Announce Type: new Abstract: Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-w…
Diagnosing Compositional Generalization in Sequential Robot Tasks
arXiv:2607.29687v1 Announce Type: new Abstract: Sequential robot manipulation requires policies to execute novel combinations of familiar instruction components…
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
arXiv:2607.29559v1 Announce Type: cross Abstract: Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward functio…
SAGP: Semantic Affordance-Guided Grasp Planning via Coarse-Zone VLM Reasoning
arXiv:2607.29374v1 Announce Type: new Abstract: Geometry-based grasp planners ensure physically valid grasps but ignore functional semantics, often generating g…
Safe Vision Language Action Models via Barrier Enhanced Flow Matching
arXiv:2607.29569v1 Announce Type: new Abstract: This article presents a modular inference framework that integrates Flow Matching generative models with formal …
Vision-Based Agile Landing on Turbulent Waters
arXiv:2605.23717v2 Announce Type: replace Abstract: Autonomous landing of Unmanned Aerial Vehicles on maritime vessels is challenging due to the coupled motion …
Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving
arXiv:2607.29052v1 Announce Type: new Abstract: End-to-end (E2E) autonomous driving aims to learn a direct mapping from visual observations to control actions. …
Receding-Horizon Next-Best-View Planner for Autonomous Leaf Surface Reconstruction
arXiv:2607.28995v1 Announce Type: new Abstract: Accurate plant leaf modeling is fundamental to downstream tasks such as plant growth monitoring, and phenotyping…
TransGraspNet: Physically and Geometrically Consistent Manipulation of Transparent Labware
arXiv:2607.29567v1 Announce Type: new Abstract: Manipulating transparent laboratory glassware that contains liquid is inherently safety-critical: even small geo…
STAGE: STyle-controllable Action GEneration for personalized autonomous driving
arXiv:2607.29517v1 Announce Type: new Abstract: Driving style refers to the behavioral preferences that drivers maintain during driving, shaped by their diverse…
MDIR: A Task-Manifold Impedance Retargeting Method for Contact-Rich Teleoperation
arXiv:2607.29271v1 Announce Type: new Abstract: Fixed Cartesian impedance makes contact-rich teleoperation demonstrations practical, but gains that secure progr…
Automated Straight-line Sewing of Stretchable Fabrics with Different Lengths
arXiv:2607.29464v1 Announce Type: new Abstract: Different Length Alignment Sewing (DLAS), which involves stretching the shorter fabric to match the longer one a…
Advances, challenges, and opportunities for legged robots
arXiv:2607.28952v1 Announce Type: new Abstract: Humanoid and quadrupedal robots have the potential to revolutionize the way we work, interact, and coexist with …