Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesFlatLab: A Unified Methodology Framework and Simulation-Based Benchmark for Robotic Manipulation of Flat Objects
arXiv:2608.14049v1 Announce Type: new Abstract: Robotic manipulation of flat objects is challenging due to the ungraspable configurations and strong variations …
hint$^2$: Hierarchical World Models for Inference-Time Temporal Logic Guidance
arXiv:2608.13678v1 Announce Type: new Abstract: A central goal of robot learning is to enable robots to execute rich instructions specified at runtime. Large-sc…
Knowledge-Data-Dual-Driven Reinforcement Learning for Autonomous Vehicle Control in Mixed Traffic
arXiv:2608.13878v1 Announce Type: new Abstract: In mixed traffic, decision-making for autonomous vehicles (AVs) confronts three interrelated challenges. First, …
BICPO-VLA: Behavior-Identified Continuation Preference Optimization for Smooth Asynchronous Vision-Language-Action Control
arXiv:2608.13924v1 Announce Type: new Abstract: The request-to-handoff gap has three coupled sources: ambiguity about the behavior intended at request time, phy…
Demonstration of Space Robot Teleoperation over a Lossy and Delayed Network using ATMOS
arXiv:2608.14031v1 Announce Type: new Abstract: We present a demonstration showcasing the Autonomy Testbed for Multi-purpose Orbiting Systems (ATMOS), a planar …
PILOT: Privileged Imitation Learning for End-to-End Motion Planning of Autonomous UAVs under Partial Observability
arXiv:2608.14082v1 Announce Type: new Abstract: Autonomous navigation in cluttered environments is hampered by partial observability and dynamic constraints. Th…
OccPlanner: Goal-Aware Occupancy-Conditioned Diffusion Planner for Pixel-Goal Navigation
arXiv:2608.14160v1 Announce Type: new Abstract: Pixel-goal navigation specifies targets directly in the agent's camera view, but a target pixel provides neither…
MMUSV-Sim: A Perception-Oriented Simulation and Data-Generation Platform for Multi-USV Cooperative Perception
arXiv:2608.14207v1 Announce Type: new Abstract: Cooperative perception among multiple unmanned surface vehicles (USVs) combines complementary observations to ex…
Accelerating Large-scale Bundle Adjustment for LiDAR Mapping via Parallel Computing
arXiv:2608.14266v1 Announce Type: new Abstract: LiDAR bundle adjustment is widely utilized in mapping to construct globally consistent point cloud maps. In this…
Control-Informed Constraint Adaptation in Minimum-Time Trajectory Planning for Autonomous Racing
arXiv:2608.14448v1 Announce Type: new Abstract: Autonomous racecars operate at the limits of vehicle dynamics, where small control errors translate into safety-…
Expected Free Energy-based Informative Path Planning for Robotic Mars Exploration
arXiv:2608.14466v1 Announce Type: new Abstract: An autonomous robot efficiently exploring an unknown environment, such as looking for water sources on Mars, fac…
Spatiotemporal Tube-Based Safety-Certificate for Autonomous Navigation of Articulated Vehicles
arXiv:2608.14531v1 Announce Type: new Abstract: Articulated vehicles are the workhorses of freight transportation, and their autonomous navigation is challengin…
Learning-Guided Sparsification of Dynamic Graphs in Robotic Exploration
arXiv:2604.16509v2 Announce Type: replace Abstract: Many robotic exploration algorithms rely on graph structures for frontier-based exploration and dynamic path…
RoboSynChallenge: Mastering Real-World Dexterity via Generalizing Synthesized Manipulation Skills
arXiv:2608.12416v1 Announce Type: new Abstract: Achieving generalizable robotic manipulation remains a central challenge in embodied intelligence. Despite rapid…
Attune: A Self-Annotation Tool for Understanding Robot Operator Attention Profiles
arXiv:2608.12650v1 Announce Type: new Abstract: Deploying robot fleets in complex, real-world environments requires human operators to supervise multiple robots…
FUSE: Active Functional Affordance Grounding through Adaptive Semantic-Geometric Evidence Acquisition
arXiv:2608.12683v1 Announce Type: new Abstract: Embodied agents must often identify and interact with objects based on their function rather than their identity…
SAP-Nav: Spatial Semantic Representation Meets Active Perception for Hierarchical Open-Vocabulary Object Navigation
arXiv:2608.12707v1 Announce Type: new Abstract: Hierarchical open-vocabulary object navigation (OVON) requires agents to follow free-form instructions that may …
Genetic Fuzzy System-Based Multi-Robot Coordination for Planetary Missions
arXiv:2608.12755v1 Announce Type: new Abstract: This paper proposes a decentralized approach for a multi-robot system (MRS) using a genetic fuzzy system to perf…
AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN
arXiv:2608.12835v1 Announce Type: new Abstract: Unmanned Aerial Vehicle Vision-Language Navigation (UAV-VLN) requires agents to follow language instructions, in…
ASPIRE-VINS: Adaptive Spline-based Visual-inertial Navigation System With Robust 3D Measurement Residuals
arXiv:2608.12840v1 Announce Type: new Abstract: Visual-inertial navigation systems estimate six-degree-of-freedom motion by fusing visual and inertial data. Mod…
BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving
arXiv:2608.12854v1 Announce Type: new Abstract: Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-en…
HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments
arXiv:2608.12860v1 Announce Type: new Abstract: Vision-Language Navigation (VLN) for humanoid robots poses challenges existing benchmarks fail to address: biped…
Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning
arXiv:2608.13026v1 Announce Type: new Abstract: Outcome-driven reinforcement learning offers a scalable way to post-train vision-language-action (VLA) policies …
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
arXiv:2608.13049v1 Announce Type: new Abstract: Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expen…
Semantic Radiance Fields as Simulators for Spatial Reasoning in Real-World Scenes
arXiv:2608.13095v1 Announce Type: new Abstract: Training and evaluating spatial reasoning in embodied agents requires diverse environments that are both geometr…
S2-HWM: Sparse Event-Structured Hierarchical World Model for Long-Horizon Surgical Robot Manipulation
arXiv:2608.13103v1 Announce Type: new Abstract: Long-horizon surgical robot manipulation is challenging because task rewards are sparse, while meaningful intera…
FAM-DQ: A Dual-Quadrotor-Based Fully Actuated Aerial Manipulator for High-Torque Interaction
arXiv:2608.13220v1 Announce Type: new Abstract: Aerial physical interaction requires aerial manipulation platforms to generate large interaction forces and torq…
Manufacturing Complex Airtight Soft Pneumatic Actuators for Soft Robotics: Process Evaluation and Optimization
arXiv:2608.13233v1 Announce Type: new Abstract: Manufacturing complex soft pneumatic actuators remains challenging because geometric fidelity, compliance, struc…
NestDex: Nested Policy Learning with Copilot Assisted Teleoperation for Dexterous Manipulation
arXiv:2608.13362v1 Announce Type: new Abstract: Dexterous manipulation promises substantially richer robot interaction with the physical world, but learning the…
FIRE-VLA: Failure-Informed Self-Evolution for Vision-Language-Action Models in Autonomous Driving
arXiv:2608.13395v1 Announce Type: new Abstract: Reinforcement learning improves autonomous-driving vision-language-action (VLA) models by evaluating trajectorie…