Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesUAV Video Deblurring via Motion-Aware Diffusion: A Path to Robust Target Detection
arXiv:2608.15259v1 Announce Type: cross Abstract: Unmanned Aerial Vehicles (UAVs) play a crucial role in various scenarios ranging from disaster response to tra…
EgoTac: In-the-wild Tactile Prediction from Egocentric Vision
arXiv:2608.15060v1 Announce Type: cross Abstract: Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot lea…
NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving
arXiv:2608.14767v1 Announce Type: cross Abstract: Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Ex…
Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade
arXiv:2608.14650v1 Announce Type: cross Abstract: Existing adaptive-inference and world-action-model systems use cheap-stage outputs or predicted futures to all…
$\tau_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
arXiv:2608.16885v1 Announce Type: new Abstract: Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them co…
When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents
arXiv:2608.16806v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-…
Neurosymbolic Embodied Agents
arXiv:2608.16794v1 Announce Type: new Abstract: Language and vision-language models generate plausible embodied plans but do not guarantee executability, as the…
Observation-Constrained Joint-Space Viewpoint Optimization for Robotic Inspection of Cylindrical Cavities
arXiv:2608.16442v1 Announce Type: new Abstract: Inspection is a core capability in many mobile robotics applications, including industrial facility monitoring, …
Planner-Conditioned Diffusion for Coordinated Multi-Agent Exploration
arXiv:2608.16229v1 Announce Type: new Abstract: Coordinated multi-agent exploration requires not only efficient individual coverage but also non-redundant cover…
Arm-Aware Guided Dexterous Grasp Generation with Arm-Agnostic Grasp Models
arXiv:2608.16351v1 Announce Type: new Abstract: Dexterous grasp generation that considers arm-related constraints is crucial in real-world scenarios involving a…
Readiness Barrier Functions: Forward-Invariant Control Authority for Overactuated Multirotor Allocation
arXiv:2608.16335v1 Announce Type: new Abstract: Allocation schemes that greedily maximize a readiness metric over the actuator fiber bundle of an overactuated m…
Marker-Constrained Pose-Graph Correction for Cross-Platform Georeferencing in GNSS-Denied Environments
arXiv:2608.16281v1 Announce Type: new Abstract: Autonomous operation in GNSS-denied environments requires heterogeneous mapping pipelines to maintain a consiste…
RoboStriker: Latent-Space Strategic Games for Autonomous Humanoid Boxing
arXiv:2608.16195v1 Announce Type: new Abstract: Achieving human-level competitive intelligence and physical agility in humanoid robots remains a profound challe…
SparkVLA: Stop-Aware Hierarchical VLA with Adaptive Action Chunking for Long-Horizon Manipulation
arXiv:2608.16172v1 Announce Type: new Abstract: At every re-observation point in a hierarchical Vision-Language-Action (VLA) system, two interface decisions mus…
SurgVIL: Scaling Surgical Robot Imitation Learning with Open-source Surgical Videos
arXiv:2608.16058v1 Announce Type: new Abstract: Learning-based surgical robot autonomy requires large-scale demonstrations with synchronized videos and robot ac…
Tabletop Pen Manipulation With a Vision-Guided 4-DoF Arm
arXiv:2608.15968v1 Announce Type: new Abstract: Low-cost four-degree-of-freedom (DoF) arms are among the most accessible robotic platforms. But they are, in the…
Rotate Disks to Reach Farther: Design and Modeling of a Novel Reconfigurable Tendon Driven Manipulator
arXiv:2608.15946v1 Announce Type: new Abstract: Rerouting the tendon path in tendon driven continuum manipulators (TDCMs) enables a broad range of deformation m…
Tactile Sim2Real without Tactile Simulation via Bottlenecked Latent Reconstruction
arXiv:2608.15897v1 Announce Type: new Abstract: Robot sensor designs, particularly tactile sensors, are highly diverse and evolve rapidly. Modeling each sensor …
Grouping Auction-Consensus Algorithm for Decentralized Task Allocation in Multi-Robot Systems
arXiv:2608.15884v1 Announce Type: new Abstract: Decentralized multi-robot task allocation (MRTA) is essential for scalable and resilient autonomous systems. The…
SCORE: Shape-Conforming Regions for Flight in Enclosed, Degraded Environments
arXiv:2608.15289v1 Announce Type: new Abstract: Autonomous UAVs enter enclosed environments such as caves and collapsed structures that confine the vehicle and …
Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning
arXiv:2608.15088v1 Announce Type: new Abstract: Human-in-the-loop (HIL) online reinforcement learning for real robots must absorb human interventions quickly wh…
MotionGS-SLAM: Event-Modulated Gaussian Splatting for Motion-Blur Robust SLAM
arXiv:2608.15024v1 Announce Type: new Abstract: Current Vision-based SLAM systems fail catastrophically when motion blur corrupts the visual input, as they atte…
HP2-SLAM: Adaptive Hybrid ICP for Robust and Efficient LiDAR SLAM
arXiv:2608.14996v1 Announce Type: new Abstract: Achieving robustness, accuracy, and efficiency simultaneously remains a central challenge in light detection and…
Evidence of Absence: Cross-Modal Abductive Risk Perception to Sustain World Models When Vision Fails
arXiv:2608.14952v1 Announce Type: new Abstract: A structured world-state (entities, relations, context, and predictive cues) is designed to preserve prediction-…
Geometry-Aware Online Mapping for 3D Gaussian Splatting SLAM
arXiv:2608.14902v1 Announce Type: new Abstract: Recent 3D Gaussian Splatting (3DGS) has enabled efficient photorealistic view synthesis and is rapidly being ado…
Modeling and Control of an Eel-Inspired Soft Robot for Design Optimization
arXiv:2608.14860v1 Announce Type: new Abstract: Anguilliform locomotion is a highly efficient swimming mode; the advent of new materials for soft robots enables…
MISTac: A Vision-Based Tactile Sensor for Minimally Invasive Surgery
arXiv:2608.14772v1 Announce Type: new Abstract: Minimally invasive and robot-assisted surgery offer many advantages over traditional open surgery, but deprive s…
SkillComposer: Learning Reusable Skills for Natural-Language Robot Programming
arXiv:2608.14944v1 Announce Type: new Abstract: Natural-language interfaces can lower the barrier to programming robots, but existing systems struggle when user…
ForceU-VLA: A Force-Aware Vision-Language-Action Model for Embodied Ultrasound Scanning
arXiv:2608.15009v1 Announce Type: new Abstract: Embodied intelligent ultrasound scanning enables the automation and standardization of the ultrasound examinatio…
PACE: Phase-Progress-Aware Credit for Long-Horizon Embodied Manipulation
arXiv:2608.15026v1 Announce Type: new Abstract: Post-training of vision-language-action (VLA) models typically relies on expert demonstrations and policy intera…