Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesHaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy
arXiv:2609.09941v1 Announce Type: new Abstract: Generalist robot policies have demonstrated strong generalization across robotic manipulation tasks, yet their s…
ViBe: Visual Behavior Adaptation for Perceptive Humanoid Whole-Body Control
arXiv:2609.09918v1 Announce Type: new Abstract: Motion tracking provides a scalable recipe for humanoid whole-body control. By design, the resulting trackers la…
PccDiffuser: Multi-solution Motion Planning for Continuum Robots
arXiv:2609.09745v1 Announce Type: new Abstract: We present the PccDiffuser, a conditional diffusion framework for continuum robots that learns a multimodal dist…
A Risk-Sensitive and Uncertainty-Aware Decision-Making and Control Framework for Safe and Robust Autonomous Driving
arXiv:2609.09650v1 Announce Type: new Abstract: Reinforcement learning (RL) has demonstrated considerable potential for autonomous driving decision-making. Howe…
JEPA Policy: Diffusion-Free Imitation Learning via Paired Action and Future Representation Prediction
arXiv:2609.09630v1 Announce Type: new Abstract: Standard behavior cloning supervises actions without explicitly constraining the future representation paired wi…
PGMT: Perceptive General Motion Tracking for Humanoid Robots
arXiv:2609.08511v2 Announce Type: replace Abstract: Humanoid motion trackers can reproduce diverse whole-body motions, but their performance degrades on complex…
Meta-RL with Bayesian Linear Task Models
arXiv:2512.20974v4 Announce Type: replace-cross Abstract: Deep Bayesian reinforcement learning adapts to unseen tasks by inferring latent transition and reward …
FiberTune: Preserving Action-Fiber Visual Residuals in Vision-Language-Action Fine-Tuning
arXiv:2606.08653v2 Announce Type: replace-cross Abstract: Action-supervised fine-tuning of vision-language-action (VLA) policies fits demonstrations effectively…
Vention opens Physical AI Lab for manufacturing in Montreal
Vention's new Physical AI Lab will focus on advancing robotic manipulation from research to scalable production-line deployment. The post Vention opens Physical…
Alignment Under Pressure: AR-HMD Support Tools for Action Teams
arXiv:2502.17295v2 Announce Type: replace-cross Abstract: Team communication breakdowns represent a contributor to patient safety risks within action teams-defi…
APEX-RBD: Mixed-Precision Exploration Framework for Hardware-Efficient Robot Dynamics Accelerator Design
arXiv:2609.05161v1 Announce Type: cross Abstract: Rigid Body Dynamics (RBD) forms the computational core of real-time robotic control, but its immense computati…
One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation
arXiv:2609.04921v1 Announce Type: cross Abstract: Diffusion probabilistic models can capture the multi-modal, interaction-rich distribution of joint future traj…
Dual-Part Multi-Lateral Branched Network for Multi-Class Segmentation in Cardiovascular Catheterization Angiograms
arXiv:2609.04590v1 Announce Type: cross Abstract: Catheterisation image processing requires segmentation models that are fast, accurate and explainable. While m…
Achieving Asymptotic Near-Optimality Without $\delta$-Similarity
arXiv:2609.04464v1 Announce Type: new Abstract: Sampling-based motion planning algorithms are a popular class of trajectory planning algorithm due to their spee…
SocioGesture: Real-Time and Adaptive Social Gesture Perception for Human-Robot Interaction
arXiv:2609.04545v1 Announce Type: new Abstract: Robots interacting with people must recognize not only explicit commands, but also social cues such as invitatio…
NavArena: Automated Construction of Goal-Oriented Navigation Benchmarks from 3D Gaussian Splatting Reconstructions
arXiv:2609.04602v1 Announce Type: new Abstract: Fixed 3D Gaussian Splatting (3DGS) reconstructions provide realistic novel views but lack the traversability con…
ToPos: Automated Optimal Positioning on Topographic Manifolds using Constrained Geodesic Voronoi Decomposition
arXiv:2609.05084v1 Announce Type: new Abstract: Reliable autonomous mapping, environmental sampling, last-mile logistics, and infrastructure deployment depend o…
Sound-based Multi-Person 3D Pose Estimation
arXiv:2609.04902v1 Announce Type: cross Abstract: Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to esti…
EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents
arXiv:2604.18271v2 Announce Type: replace Abstract: As the world of agentic artificial intelligence applied to robotics evolves, the need for agents capable of …
Toward Context-Aware Exoskeleton Assistance: Integrating Computer Vision Payload Estimation with a Multi-Metric Optimization Space
arXiv:2508.06207v3 Announce Type: replace Abstract: Back-support exoskeletons mitigate musculoskeletal strain, yet current systems rely on reactive sensing and …
Where Appearance Fails, Geometry Recognizes: A CAD-Free 3D Shape Prior That Complements Vision Foundation Models
arXiv:2609.04381v1 Announce Type: cross Abstract: Recognizing specific objects onboarded without a labeled training set recurs across manufacturing and service …
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
arXiv:2609.05324v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. …
Morphology and actuation as inductive biases in robotic hand manipulation
arXiv:2609.05206v1 Announce Type: new Abstract: Robotic hands vary widely in anatomical fidelity and mechanical complexity, and these structural choices influen…
LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models
arXiv:2609.05178v1 Announce Type: new Abstract: Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in r…
A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning
arXiv:2609.05133v1 Announce Type: new Abstract: This paper addresses navigation by composite heterogeneous robots in a decentralized system when policy reasonin…
HaptiNet: Networked Haptic Robots Enable Physical Co-presence in Geographically-Unconstrained Rehabilitation
arXiv:2609.04799v1 Announce Type: new Abstract: Cooperative rehabilitation enhances engagement, task performance, and social-motor interaction, yet it demands p…
Continuous Cognitive Coverage for Autonomous Robots via Event-Dependent Cognitive Treatment and Learning
arXiv:2609.04770v1 Announce Type: new Abstract: Autonomous robots continuously encounter objects, changes, and situations, and every event admitted into cogniti…
AquaBEV: Monocular Underwater BEV Occupancy with 3D Sonar Supervision
arXiv:2609.04411v1 Announce Type: new Abstract: Autonomous underwater robots are widely used for exploration, monitoring, and inspection, where safe navigation …
MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision
arXiv:2609.04958v1 Announce Type: cross Abstract: Recovering camera and hand motion in world coordinates from egocentric video is a key capability for activity …
CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Surround Depth Estimation
arXiv:2609.05397v1 Announce Type: cross Abstract: Reliable 3D understanding of the surrounding environment is a core requirement for autonomous driving. Multi-v…