Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesFastTrack: GPU-Accelerated Tracking for Visual SLAM
arXiv:2509.10757v2 Announce Type: replace Abstract: The tracking module of a visual-inertial SLAM system processes incoming image frames and IMU data to estimat…
ZAPS-DA: Zero-Phase Action Policy Smoothing with Decoupled Actor for Continuous Control in Reinforcement Learning
arXiv:2605.30612v2 Announce Type: replace Abstract: Continuous control policies trained with off-policy reinforcement learning frequently exhibit high-frequency…
BiNoMaP: Learning Category-Level Bimanual Non-Prehensile Manipulation Primitives
arXiv:2509.21256v3 Announce Type: replace Abstract: Non-prehensile manipulation, encompassing ungraspable actions such as pushing, poking, pivoting, and wrappin…
TFP: Temporally Conditioned Memory-Fusion Policies for Visuomotor Learning
arXiv:2607.08283v1 Announce Type: new Abstract: Vision--Language--Action (VLA) policies such as $\pi_{0.5}$ and OpenVLA perform well on many manipulation tasks,…
OREN: Octree Residual Network for Real-Time Euclidean Signed Distance Mapping
arXiv:2510.18999v3 Announce Type: replace Abstract: Reconstructing signed distance functions (SDFs) from point cloud data benefits many robot autonomy capabilit…
SkillPlug: Unsupervised Skill Mining for Few-Shot Adaptation in Robotic Manipulation
arXiv:2607.08354v1 Announce Type: new Abstract: Learning transferable visuomotor imitation policies that generalize across diverse manipulation tasks and adapt …
Factors Influencing Conversational Engagement in Robot-Delivered Individual Cognitive Stimulation Therapy (iCST) for Dementia in Home Settings
arXiv:2607.07998v1 Announce Type: new Abstract: Social robots offer a promising means of supporting cognitive therapies for dementia care by guiding structured …
EVIS: A Physics-Grounded Event Camera Plugin for NVIDIA Isaac Sim
arXiv:2607.08098v1 Announce Type: cross Abstract: Event cameras offer microsecond temporal resolution, low latency, and high dynamic range, making them attracti…
INTENT: An LSTM Framework for Vehicle Intention Prediction in Intersection Scenarios with Comprehensive Ablation Analysis
arXiv:2607.08316v1 Announce Type: cross Abstract: Vehicle intention prediction is a pivotal aspect in the agility and safety of autonomous vehicles in all drivi…
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
arXiv:2507.05240v2 Announce Type: replace Abstract: Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual str…
V-VLAPS: Value-Guided Planning for Vision-Language-Action Models
arXiv:2601.00969v3 Announce Type: replace Abstract: Vision-language-action (VLA) models provide strong action priors for robotic manipulation, but their reactiv…
APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts
arXiv:2607.08024v1 Announce Type: cross Abstract: Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility.…
Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference
arXiv:2607.08724v1 Announce Type: cross Abstract: Human decision-making is highly flexible -- some actions are taken immediately; others require longer delibera…
TriphiBot: A Triphibious Robot Combining FOC-based Propulsion with Eccentric Design
arXiv:2602.01385v2 Announce Type: replace Abstract: Triphibious robots capable of multi-domain motion and cross-domain transitions are promising to handle compl…
FunHOI: Annotation-Free 3D Hand-Object Interaction Generation via Functional Text Guidance
arXiv:2502.20805v3 Announce Type: replace Abstract: Hand-object interaction(HOI) is the fundamental link between human and environment, yet its dexterous and co…
Input-Constrained Spatiotemporal Tubes for Safe Navigation of Unknown Euler-Lagrange Systems in Dynamic Environments
arXiv:2607.08189v1 Announce Type: cross Abstract: Safe navigation in dynamic environments is challenging when system dynamics are unknown and actuator inputs ar…
AnyDexRT: Calibration-Free Dexterous Hand Retargeting with Few-Shot Human Guidance
arXiv:2607.08341v1 Announce Type: new Abstract: Teleoperation is a key interface for controlling dexterous robotic hands and collecting demonstrations for imita…
A Collaborative Reasoning Framework for Anomaly Diagnostics in Underwater Robotics
arXiv:2511.03075v2 Announce Type: replace Abstract: The safe deployment of autonomous systems in safety-critical settings requires a paradigm that combines huma…
Design optimization and robustness analysis of rigid-link flapping mechanisms
arXiv:2503.21204v3 Announce Type: replace Abstract: Rigid link flapping mechanisms remain the most practical choice for flapping wing micro-aerial vehicles (MAV…
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
arXiv:2607.08741v1 Announce Type: cross Abstract: Generating realistic 3D human motions in real-time within interactive applications is key for animation, simul…
Native Video-Action Pretraining for Generalizable Robot Control
arXiv:2607.08639v1 Announce Type: new Abstract: The advent of video-action models offers a promising path for robot control. Nevertheless, we argue that repurpo…
Towards Soft Robotic Exogloves for Musculoskeletal Manipulation to Reduce Pain and Spasticity
arXiv:2607.07958v1 Announce Type: new Abstract: Hand spasticity and resulting pain affect 12 million people worldwide, including stroke survivors, arthritis pat…
EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data
arXiv:2607.08436v1 Announce Type: new Abstract: Egocentric human data offers scalable supervision for robot manipulation. However, behavior cloning entangles tr…
Early to Share, Late to Save: Synchronisation-Driven Communication Gating in Bandwidth-Constrained Cooperative VLN
arXiv:2607.08504v1 Announce Type: cross Abstract: Most cooperative Vision-Language Navigation (VLN) methods assume unlimited communication, not considering real…
SASGeo: Stability-Aware Semantic Map Localization for GNSS-Denied UAVs -- A Framework and Synthetic Proof of Concept
arXiv:2607.07737v1 Announce Type: new Abstract: GNSS-denied unmanned aerial vehicles require occasional absolute position fixes to bound the drift of visual-ine…
Monocular Vision Based Control Framework for Grasping
arXiv:2607.07897v1 Announce Type: new Abstract: Grasping in unstructured environments requires handling objects with widely different mechanical properties, fro…
In vivo feasibility study of humanoid robots in surgery
arXiv:2607.07972v1 Announce Type: new Abstract: Recent advances in actuation, control and learning have rapidly pushed humanoid robots from a distant vision tow…
ContactMimic: Humanoid Object Interaction via Contact Control
arXiv:2607.08742v1 Announce Type: new Abstract: Keypoint tracking alone is insufficient for object interaction tasks such as sitting on a chair, wiping a board,…
Idiobionics: The Unification of Privacy and Intelligent Robotic Prostheses
arXiv:2607.07775v1 Announce Type: cross Abstract: The human body is at the center of a growing family of technologies designed to tightly and persistently coupl…
SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment
arXiv:2511.08583v2 Announce Type: replace Abstract: Developing efficient and accurate visuomotor policies poses a central challenge in robotic imitation learnin…