Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2771 storiesToken-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation
arXiv:2607.16806v1 Announce Type: new Abstract: Vision-Language Navigation in dynamic, human-centric environments exposes a fundamental tension: linguistic reas…
The World According to a Social Robot -- Augmenting Human-Robot Dialogue With Vision Language Models
arXiv:2607.16318v1 Announce Type: cross Abstract: Vision Language Models (VLMs) enable robots to visually perceive their environment as well as the actions and …
TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies
arXiv:2606.06491v2 Announce Type: replace Abstract: Robot manipulation alternates between low-risk transit phases that call for fast execution and high-risk con…
Li-ViP3D++: Query-Gated Deformable Camera-LiDAR Fusion for End-to-End Perception and Trajectory Prediction
arXiv:2601.20720v2 Announce Type: replace-cross Abstract: End-to-end perception and trajectory prediction from raw sensor data is one of the key capabilities fo…
Differentiable Reinforcement Learning for Path Tracking by an Agile Fish-Like Robot
arXiv:2607.16508v1 Announce Type: new Abstract: Fish-like swimming has inspired the design of several dozens if not hundreds of bioinspired robots in the last f…
A2RL V\textsubscript{max}: The A2RL autonomous racing dataset for long-range, high-speed perception and multi-vehicle interaction
arXiv:2607.17813v1 Announce Type: new Abstract: In autonomous driving development, a perception dataset is crucial, as it provides fundamental data for training…
Lifelong Multi-Subsystem Pickup and Delivery with Buffer-Limited Handover Stations
arXiv:2607.17724v1 Announce Type: new Abstract: Coordinating payload transfers between subsystems is a critical challenge in lifelong Multi-Agent Pickup and Del…
DROID-ANCHOR: Odometry-Anchored Recurrent Metric Depth Estimation
arXiv:2607.17058v1 Announce Type: new Abstract: Precise metric depth estimation is fundamental for autonomous robot navigation, yet monocular systems inherently…
GhostShell: Streaming LLM Function Calls for Concurrent Embodied Programming
arXiv:2508.05298v3 Announce Type: replace Abstract: We present GhostShell, a novel approach that leverages Large Language Models (LLMs) for streaming and concur…
S.E.A.G.R: A Socially and Emotionally Aware Greeting Robot Framework with Dual-Layer Cultural and Affective Modulation
arXiv:2607.16341v1 Announce Type: new Abstract: This paper presents SEAGR (Socially and Emotionally Aware Greeting Robot), a robotic greeting framework designed…
Lifelong Localization in Dynamic Indoor Environments Combining Odometry with Sparse Distance Sampling
arXiv:2607.17852v1 Announce Type: new Abstract: Localization is a key task in robot navigation, and many techniques exist for it. In many plausible scenarios, a…
SplatSearch: Instance Image Goal Navigation for Mobile Robots using 3D Gaussian Splatting and Diffusion Models
arXiv:2511.12972v2 Announce Type: replace Abstract: The Instance Image Goal Navigation (IIN) problem requires mobile robots deployed in unknown environments to …
Technical Design Review of Duke Robotics Club's Oogway & Crush: AUVs for RoboSub 2026
arXiv:2607.18075v1 Announce Type: new Abstract: The Duke Robotics Club presents Oogway and Crush, our AUVs for RoboSub 2026. This year's strategy expands on our…
PREFAIL: Identifying Precursors to Failures in Robotic Lift-and-Place Tasks to Improve Task Execution Performance
arXiv:2607.16921v1 Announce Type: new Abstract: Non-prehensile manipulation enables flexible material handling with part carriers, but friction-based support ma…
HyperDCM: Dynamic Cluster Memory Replay in Hyperbolic Space for Continual Robotic Navigation Across Scenes
arXiv:2607.16267v1 Announce Type: new Abstract: Continual learning in visual navigation remains challenging due to catastrophic forgetting and the difficulties …
Receiver-Centered Robot-to-Human Handover with Grasp-Aware Object Orientation
arXiv:2607.17839v1 Announce Type: new Abstract: Collaborative robots are increasingly sharing workspaces with human operators, making tool handover a frequent a…
Rapid and Safe Trajectory Planning over Diverse Scenes through Diffusion Composition
arXiv:2507.04384v4 Announce Type: replace Abstract: Achieving safe, efficient, and kinematically feasible planning in dynamic environments remains a significant…
Linear Stability Analysis of an INDI Pitch-Rate Controller under Model Mismatch for a Tilt-Rotor VTOL UAV
arXiv:2607.16471v1 Announce Type: new Abstract: Incremental Nonlinear Dynamic Inversion (INDI) is attractive for unmanned aerial vehicle (UAV) flight control be…
Does Robust VIO Need More Learning? Geometry-Verified Visual Measurements under Distribution Shift
arXiv:2607.17956v1 Announce Type: new Abstract: Learning is increasingly introduced into visual-inertial odometry (VIO), ranging from learned feature front-ends…
Task-Space Constrained Stochastic Trajectory Optimization for Time-Optimal Forestry Crane Motion Planning
arXiv:2607.17818v1 Announce Type: new Abstract: Efficient, collision-free, and time-optimal motion planning is a fundamental requirement for autonomous forestry…
Closing the Loop in Humanoid VLA: Persistent 3D Object Tokens for Verifiable Loco-Manipulation
arXiv:2607.18016v1 Announce Type: new Abstract: Vision-language-action policies are a promising foundation for general robot control, but long-horizon humanoid …
Hybrid Machine Learning for Articulation Angle Estimation of Truck-Semitrailer Combinations
arXiv:2607.16758v1 Announce Type: cross Abstract: Accurate articulation angle estimation of trucks with trailers is critical for autonomous driving and advanced…
Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering
arXiv:2510.01483v3 Announce Type: replace Abstract: Vision-language models (VLMs) demonstrate strong image-level scene understanding, but reasoning over long eg…
DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks
arXiv:2511.14592v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-cr…
Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning
arXiv:2607.17760v1 Announce Type: cross Abstract: Inverse reinforcement learning (IRL) provides a powerful framework for learning from demonstrations. However, …
Real-Time sEMG-Based Telecontrol of an Assistive Robotic Arm Using a 1D Convolutional Neural Network
arXiv:2607.16310v1 Announce Type: new Abstract: Motor impairments affecting the upper limb significantly reduce autonomy in daily activities, particularly for t…
Leveraging Two Robotic Arms for Tight Assembly Performance Gains
arXiv:2607.17876v1 Announce Type: new Abstract: We provide a novel end-to-end framework for the execution of an assembly operation by two robotic arms, given th…
MEVION: Low-Cost Open-Source Data Collection System for Powerful and High-Speed Dual-Arm Manipulation
arXiv:2607.17970v1 Announce Type: new Abstract: The global competition for developing robotic foundation models is intensifying. Among the data collection syste…
Polar Coordinate-based Differential Evolution for Moving Target Search Using Vision Sensor on Unmanned Aerial Vehicles
arXiv:2607.17771v1 Announce Type: new Abstract: In search and rescue operations, there is a period known as the "golden time" during which the probability of fi…
Test-Time Scaling for World Action Models via Zero-Shot Geometric Evaluation
arXiv:2607.17454v1 Announce Type: new Abstract: Test-time scaling improves foundation-model inference by spending additional computation, but robot control requ…