Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesTouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation
arXiv:2607.07287v1 Announce Type: new Abstract: Dexterous manipulation in everyday environments requires both anticipation and reaction: a robot must predict ho…
Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum
arXiv:2605.21133v2 Announce Type: replace Abstract: In this paper, we explore spatial-aware humanoid whole-body manipulation task. Compared with tabletop settin…
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
arXiv:2607.06988v1 Announce Type: new Abstract: Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging…
GeoGS-SLAM: Geometry-Only Gaussian Splatting for Dense Monocular SLAM
arXiv:2607.07452v1 Announce Type: new Abstract: Dense visual SLAM is a fundamental problem in robotics. Recent advances in 3DGS have demonstrated its potential …
End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent
arXiv:2607.06964v1 Announce Type: new Abstract: Bridging the gap between human pilot intent and autonomous flight operation is critical for real-world electric …
MiLSD: A Micro Line-Segment Detector for Resource-Constrained Devices
arXiv:2607.06600v1 Announce Type: cross Abstract: Line segment detection is a key building block in visual SLAM, 3D reconstruction, and industrial inspection. R…
How the Fusion of Onboard Sensors and V2X Data can Improve (or not) the Cooperative Perception of Connected Automated Vehicles
arXiv:2607.07114v1 Announce Type: cross Abstract: Automated vehicles rely on onboard sensors to perceive their surroundings and navigate autonomously. However, …
Generating Personalized Lower-Limb Kinematics Across Walking Speeds Using Subject-Conditioned Diffusion
arXiv:2607.07533v1 Announce Type: new Abstract: Personalizing exoskeleton assistance requires user-specific gait data across many locomotor tasks, yet collectin…
Continuous and large-scale: ELEANOR, the soft architected arm inspired by the elephant trunk
arXiv:2607.07622v1 Announce Type: new Abstract: The elephant trunk is a dexterous and versatile manipulator whose performance is still unmatched in robotics. In…
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
arXiv:2507.05116v5 Announce Type: replace-cross Abstract: Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic mani…
Communicative Efficiency of Single vs. Multi-Axis Robot Neck Motion
arXiv:2607.07390v1 Announce Type: new Abstract: Nonverbal communication through head and neck movement is fundamental to human social signalling, yet how roboti…
Momentum Based Reward Design for Low Emission Traffic Signal Control
arXiv:2605.29693v2 Announce Type: replace-cross Abstract: Urban traffic congestion is a growing global issue contributing significantly to long commute times an…
GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model
arXiv:2607.06882v1 Announce Type: new Abstract: Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated vi…
Can We Really Learn One Representation to Optimize All Rewards?
arXiv:2602.11399v2 Announce Type: replace-cross Abstract: As unsupervised pretraining becomes increasingly ubiquitous in reinforcement learning, a more thorough…
Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation
arXiv:2607.06957v1 Announce Type: new Abstract: Realistic and diverse traffic simulation is essential to autonomous driving development. Yet prevailing benchmar…
A Continual Learning Framework for Adaptive Control of Modular Soft Robots
arXiv:2607.06740v1 Announce Type: new Abstract: Soft robots have attracted significant attention in applications such as medical intervention, rehabilitation, a…
PriGo: Test-Time Primitive Guidance to Diffusion and Flow Policies for Adaptive Robotic Manipulation
arXiv:2607.07076v1 Announce Type: new Abstract: Imitation learning has enabled remarkable progress in robotic manipulation, especially with diffusion and flow-b…
SonoRank: Towards Calibration-Free Real-Time Finger Flexion Detection from Forearm Ultrasound Sequences
arXiv:2607.07542v1 Announce Type: new Abstract: Powered prosthetic hands are frequently abandoned, largely due to the limited functionality of current devices t…
CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis
arXiv:2607.07601v1 Announce Type: new Abstract: Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulato…
Pelican-VLA 0.5: Attending Before Acting Benefits Generalization
arXiv:2607.06655v1 Announce Type: new Abstract: In this report, we present Pelican-VLA 0.5, a unified VLA model that integrates vision-language understanding, f…
Manual, Joystick, or Haptic Control? An In Vitro Comparison of Navigation Strategies for Robotic Interventional Neuroradiology Procedures
arXiv:2607.07253v1 Announce Type: new Abstract: Objective: To evaluate robotic controller interfaces for interventional neuroradiology procedures in-vitro incor…
Agent-Exploitation Affordances: From Basic to Complex Representation Patterns
arXiv:2607.07475v1 Announce Type: new Abstract: In robotics, the capability of an artificial agent to represent the range of its action possibilities, i.e. affo…
PLED-VINS: A Point-Line Event-Based Visual Inertial SLAM for Dynamic Environments
arXiv:2607.07374v1 Announce Type: new Abstract: Dynamic environments remain a fundamental challenge for visual SLAM, where unreliable observations from moving o…
DAG-Based QoS-Aware Dynamic Task Placement for Networked Multi-Stage Control Pipelines
arXiv:2605.19887v2 Announce Type: replace-cross Abstract: Current Physical AI (PAI) relies heavily on closed-loop visual-servoing pipelines, whose perception an…
Object Search in Partially-Known Environments via LLM-informed Model-based Planning and Prompt Selection
arXiv:2603.23800v2 Announce Type: replace Abstract: We present a novel LLM-informed model-based planning framework, and a novel prompt selection method, for obj…
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
arXiv:2607.07608v1 Announce Type: new Abstract: Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Ma…
Smooth Operator: A Real-Time Sampling-Based Algorithm for Kinematic Hand Retargeting
arXiv:2607.07491v1 Announce Type: new Abstract: Advances in learning-based robotic manipulation, such as Vision-Language-Action (VLA) models and Video Action Mo…
Safe Reinforcement Learning using Ideas from Model Predictive Control
arXiv:2607.07252v1 Announce Type: cross Abstract: Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly app…
Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer
arXiv:2510.24108v2 Announce Type: replace Abstract: Human demonstrations are widely considered the cornerstone of end-to-end (E2E) autonomous driving despite hu…
EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI
arXiv:2607.07459v1 Announce Type: new Abstract: We present EmbodiedGen V2, a generative 3D world engine for building executable sim-ready environments for embod…