Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesStipple: Real-Time Incremental Gaussian Splatting with Visual-Inertial Tracking
arXiv:2608.00931v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) provides efficient rendering of photo-realistic scenes, but its heavy preprocessing…
TWINS: A Tactile Wearable Isomorphic Arm Networked System for Contact-Rich Manipulation Learning
arXiv:2608.01733v1 Announce Type: new Abstract: Recent advances in robot learning for manipulation have increased the importance of collecting real-world demons…
First Deployable Dynamic-CoM: A Unified Policy and Method-Agnostic Benchmark for Humanoid Single-Leg Balance
arXiv:2608.00500v1 Announce Type: new Abstract: Unified humanoid policies handle agile whole-body motion, yet stumble on a simple demand: staying balanced on on…
RobotDancing: Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking
arXiv:2509.20717v2 Announce Type: replace Abstract: Long-horizon, high-dynamic motion tracking on humanoids remains brittle: retargeted reference motions are ty…
Certifying Plans under Model Mismatch: A Trilemma for Reachability from Scarce Data
arXiv:2608.02453v1 Announce Type: new Abstract: Sim-to-real policies are designed under nominal dynamics, but target-system trials may yield only a few isolated…
Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?
arXiv:2608.02547v1 Announce Type: new Abstract: Action chunking---predicting and executing multiple actions instead of a single action---has proven to be a crit…
Assistant Placement Aria: A Benchmark for Egocentric Placement Assistance
arXiv:2608.00652v1 Announce Type: new Abstract: Human assistance in robotics spans around several tasks such as navigation, object manipulation, and placement, …
ChainVLA: Chaining Vision-Language-Action Queries through a Unified Execution State for Long-Horizon Manipulation
arXiv:2608.02326v1 Announce Type: new Abstract: Humans perform long-horizon manipulation by retaining knowledge of what earlier actions have established while c…
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
arXiv:2608.02580v1 Announce Type: new Abstract: Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentr…
Action Chunk Scheduling for Batched Robot Policy Serving
arXiv:2608.00337v1 Announce Type: new Abstract: Deploying robot foundation models at scale is the next step towards realizing the potential of general-purpose r…
Embodied Passive Aeroacoustic Perception Enables Relative Sensing and Pursuit Between Aerial Robots
arXiv:2608.00401v1 Announce Type: new Abstract: Aerial robots generate structured aeroacoustic fields during flight, yet these signals have been underexplored a…
Developing Combined Manipulation and Locomotion Skills with Interaction Representation and Skill Composition
arXiv:2608.00208v1 Announce Type: new Abstract: This paper addresses how to enable a humanoid robot to learn motion policies based on developmental principles a…
Compliant Sphere Lattice Contact: Distributed Contact Modeling for Sphere-Based Robot Representations
arXiv:2608.00263v1 Announce Type: new Abstract: Contact planning in robotics requires models that are both computationally efficient and physically accurate. Sp…
OmniAI: A Surface-Adaptive Aerial Projection Interface for Human--Drone Interaction
arXiv:2608.00721v1 Announce Type: new Abstract: Drones in human environments often lack spatially grounded in- terfaces for situated communication. We present O…
MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving
arXiv:2608.02449v1 Announce Type: cross Abstract: Deploying vision-language models (VLMs) for safety-critical spatial reasoning on resource-constrained autonomo…
World Action Models in Real Time: An Empirical Study of Smooth Execution via Asynchronous Deployment
arXiv:2608.01880v1 Announce Type: new Abstract: World Action Models generate fixed-horizon action chunks through iterative denoising, creating substantial infer…
GraRe: Grasp Candidate Re-Ranking for Frozen 6-DoF Grasp Detectors
arXiv:2608.00946v1 Announce Type: new Abstract: Existing 6-DoF grasp detectors typically rank grasp candidates by detector confidence. However, our analysis on …
Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models
arXiv:2608.02197v1 Announce Type: new Abstract: Visual representations of VLA models remain unreliable for spatially precise robotic manipulation. We uncover th…
Uncertainty Quantification for Visual Object Pose Estimation: S-Lemma Ellipsoidal Bounds
arXiv:2511.21666v2 Announce Type: replace Abstract: Quantifying the uncertainty of an object's pose estimate is essential for robust control and planning. Altho…
Local-Canonicalization Equivariant Graph Neural Networks for Sample-Efficient and Generalizable Swarm Robot Control
arXiv:2509.14431v2 Announce Type: replace Abstract: Multi-agent reinforcement learning (MARL) policies for swarm control often learn inefficiently and generaliz…
KING: Embodiment-Aware Kinematic Graph Neural Network for Unified Motion Representation of Legged and Wheeled Robots
arXiv:2608.01015v1 Announce Type: new Abstract: Kinematic models provide reliable motion constraints for odometry estimation in featureless environments, where …
Spline Policy: A Structured Representation for Robot Policies
arXiv:2606.07386v2 Announce Type: replace Abstract: Modern imitation-learning policies for robot manipulation often represent actions as fixed-resolution action…
Staged Multi-Agent Training (SMAT) for Hip Exoskeletons: Metabolic and Biomechanical Validation of a Simulation-Trained Co-Adaptive Controller
arXiv:2608.00715v1 Announce Type: new Abstract: Learning-based controllers can deliver exoskeleton assistance after training entirely in physics-based simulatio…
Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving
arXiv:2608.00237v1 Announce Type: cross Abstract: Vision-language models (VLMs) have recently emerged as a promising paradigm for end-to-end autonomous driving,…
Altitude-Adaptive Vision-Only Geo-Localization for UAVs in GPS-Denied Environments
arXiv:2602.23872v4 Announce Type: replace-cross Abstract: Matching downward-looking unmanned aerial vehicle (UAV) images to georeferenced satellite or aerial ma…
AffordTrajDP: Dynamic Affordance-Guided Visuomotor Policy Learning for Robotic Manipulation
arXiv:2608.01603v1 Announce Type: new Abstract: Affordance-guided imitation learning has shown impressive performance in robotic manipulation tasks by compressi…
PRISM: Privileged Probabilistic Latent Supervision for End-to-End Autonomous Driving Motion Planning
arXiv:2608.01201v1 Announce Type: new Abstract: End-to-end autonomous driving (E2E AD) systems integrate perception, prediction, and planning into a single diff…
From Failures to Supervision: DynamicEnvPlan for Robust Long-Horizon Embodied Planning
arXiv:2608.00613v1 Announce Type: new Abstract: Physical-world interaction is inherently dynamic, as environments can evolve during execution, requiring agents …
ORCESTRA: VLM-driven Visual Robot programming in Mixed Reality
arXiv:2608.00775v1 Announce Type: new Abstract: ORCESTRA is a mixed-reality system for programming robot digital twins through no-code waypoint teaching and lan…
HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing
arXiv:2603.15257v2 Announce Type: replace Abstract: Tactile sensing is a crucial capability for Vision-Language-Action (VLA) architectures, as it enables dexter…