Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesToken-Based Affordance Grounding with Large Vision-Language Models
arXiv:2607.03595v1 Announce Type: cross Abstract: Affordance grounding aims to localize image regions that support a specific action, serving as a core capabili…
Nano-U: Efficient Terrain Segmentation for Tiny Robot Navigation
arXiv:2605.10210v2 Announce Type: replace Abstract: Terrain segmentation is a fundamental capability for autonomous mobile robots operating in unstructured outd…
Multi-Robot Open Adaptive Teaming Across Unseen Environments, Partners, and Scales
arXiv:2607.04972v1 Announce Type: new Abstract: Deploying robot teams in the real world requires simultaneous adaptation to unseen environments, unknown partner…
CN-CBF: Composite Neural Control Barrier Function for Robot Navigation in Dynamic Environments
arXiv:2603.06921v2 Announce Type: replace Abstract: Safe navigation of autonomous robots remains one of the core challenges in the field, especially in dynamic …
Integrating Physics-Informed Neural Networks for Safe Reinforcement Learning in a 1-DoF Helicopter System
arXiv:2607.03125v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) offers powerful control for industrial cyber-physical systems (ICPSs), but i…
Motion Attribution for Video Generation
arXiv:2601.08828v2 Announce Type: replace-cross Abstract: Despite the rapid progress of video generation models, the role of data in influencing motion is poorl…
Governed Caste Reassignment in Heterogeneous Swarms: An Asymmetric-Trust Protocol with Audited Operator Countersignature
arXiv:2607.04634v1 Announce Type: new Abstract: In heterogeneous robot swarms, caste reassignment (rebinding a robot to a new capability-bound role) is a high-f…
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
arXiv:2607.04681v1 Announce Type: new Abstract: Embodied Chain-of-Thought has emerged as a promising mechanism to enhance robot decision-making and interpretabi…
ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning
arXiv:2607.04162v1 Announce Type: new Abstract: Open-ended tabletop manipulation requires agents to not only understand natural language but also adapt to dynam…
LOTUSim: Multi-Domain Simulator for Marine Robotics
arXiv:2607.03072v1 Announce Type: cross Abstract: Simulation is essential for maritime robotics, supporting operator training, mission rehearsal, and human-vehi…
Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models
arXiv:2607.04546v1 Announce Type: new Abstract: Action-conditioned world models allow robots to predict the future consequences of candidate actions without add…
A Spiking Sequence Generator for Polar Trajectories on Neuromorphic Hardware
arXiv:2607.02753v1 Announce Type: cross Abstract: Neuromorphic controllers for size, weight, and power-constrained systems require neural architectures that are…
iVISION-2DCD: A Long-Term Change Detection Dataset for Large-Scale Outdoor Construction Monitoring
arXiv:2607.03553v1 Announce Type: cross Abstract: Automation in construction is essential for reducing costs and human errors in large-scale projects. We approa…
SilvaScenes: Tree Detection and Species Classification from Under-Canopy Images in Natural Forests
arXiv:2510.09458v2 Announce Type: replace-cross Abstract: Interest in forestry automation is growing alongside rapid advances in deep learning. In particular, t…
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
arXiv:2607.05390v1 Announce Type: new Abstract: Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and model…
Designing Touch for Trauma-Informed Social Robots: A Design Space for Direct and Indirect Actuation
arXiv:2607.04981v1 Announce Type: new Abstract: Touch is a fundamental communication modality in human-robot interaction and may support grounding, emotional re…
Robustness Verification of an Autonomous Underwater Vehicle-based Plankton Classifier
arXiv:2607.04453v1 Announce Type: new Abstract: The assessment of planktonic standing stocks and microorganism structures is critical for understanding upper oc…
Toward Interaction Dynamics: A Predictive Framework for Safe Physical Human Robot Interaction
arXiv:2606.08281v2 Announce Type: replace Abstract: Safe physical human-robot interaction (pHRI) is fundamentally a problem of interaction dynamics: the robot m…
From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model
arXiv:2607.05396v1 Announce Type: cross Abstract: Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience r…
karl. -- A Research Vehicle for Automated and Connected Driving
arXiv:2602.08842v3 Announce Type: replace-cross Abstract: As highly automated driving is transitioning from single-vehicle closed-access testing to commercial d…
Insect-inspired Visual Point-goal Navigation
arXiv:2601.16806v4 Announce Type: replace-cross Abstract: Insect neuroethology provides a compelling biological template for efficient autonomous navigation. We…
Differential Amplifier-Inspired AmpAttention for Multi-View Robotic Manipulation
arXiv:2607.02845v1 Announce Type: new Abstract: Multi-view robotic manipulation methods with the attention mechanism have recently achieved significant progress…
Anticipatory Reinforcement Learning for Trajectory Tracking
arXiv:2607.03132v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) in industrial control often suffers from lag and overshoot due to purely rea…
High-Fidelity One-Step Generative Visuomotor Policy via Recursive Correction, Frequency Consistency, and Contrastive Flow Matching
arXiv:2607.03865v1 Announce Type: new Abstract: Generative models such as diffusion and flow matching have advanced robotic visuomotor policies by modeling mult…
Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control
arXiv:2607.04978v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control …
MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments
arXiv:2512.22867v2 Announce Type: replace-cross Abstract: Socially compliant navigation requires structured reasoning about dynamic pedestrians and physical con…
DRBA: Dynamic Robotic Balance Assistant -- An assist-as-needed gait and balance rehabilitation robot for versatile training
arXiv:2607.03027v1 Announce Type: new Abstract: The decline of human balance control due to aging and pathological conditions increases fall risk, a major conce…
Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies
arXiv:2602.06575v2 Announce Type: replace Abstract: Vision-language-action (VLA) models typically inject proprioception only as a late conditioning signal, prev…
A Co-Design Framework for High-Performance Jumping of a Five-Bar Monoped with Actuator Optimization
arXiv:2604.06025v2 Announce Type: replace Abstract: The performance of legged robots depends strongly on both mechanical design and control, motivating co-desig…
Layout-independent actuation allocator for fin-actuated marine robots
arXiv:2607.03204v1 Announce Type: new Abstract: In this study, we propose a layout-independent control allocator capable of zero-shot deployment across diverse …