Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
2826 storiesCorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors
arXiv:2604.21241v2 Announce Type: replace Abstract: Vision--Language--Action (VLA) models often use intermediate representations to connect multimodal inputs wi…
Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies
arXiv:2607.05122v1 Announce Type: cross Abstract: Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain…
Toward Personalized Social Robots for Child Well-being: Data Requirement Principles from a Recommender-System Perspective
arXiv:2607.05110v1 Announce Type: cross Abstract: Social robots are increasingly deployed in clinical settings to support the well-being of children, where effe…
MOSAIC: Skill-Centric Manipulation Planning with Physics Simulation
arXiv:2504.16738v3 Announce Type: replace Abstract: Planning long-horizon manipulation motions using a set of predefined skills is a central challenge in roboti…
Finite Reliability Representations: Noise-Calibrated Belief-Space Covers for Reliable Decision-Making
arXiv:2607.04019v1 Announce Type: cross Abstract: Physical sensing and actuation noise floors should inform how much belief resolution a decision-making system …
Agent-driven Long-tail Simulation for Autonomous Driving
arXiv:2607.04331v1 Announce Type: new Abstract: Evaluating autonomous driving systems in closed-loop settings requires realistic and interactive simulation, yet…
Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation
arXiv:2607.05377v1 Announce Type: new Abstract: While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they stru…
Athena-WBC: Capability-Aligned Policy Experts for Long-Tail Humanoid Whole-Body Control
arXiv:2607.04837v1 Announce Type: new Abstract: Large-scale humanoid motion-tracking controllers are commonly improved by reallocating training effort: difficul…
HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control
arXiv:2607.03449v1 Announce Type: new Abstract: Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian …
AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning
arXiv:2607.03182v1 Announce Type: new Abstract: Autonomous driving planning requires translating navigation intent, traffic rules, dynamic interactions, and lan…
TreeLoc++: Robust 6-DoF LiDAR Localization in Forests with a Compact Digital Forest Inventory
arXiv:2603.03695v2 Announce Type: replace Abstract: Reliable localization is essential for sustainable forest management, as it allows robots to revisit and mon…
SurgAM: Surgical Affordance Map Prediction with Multimodal Feature Fusion for Robot Autonomy
arXiv:2607.04378v1 Announce Type: new Abstract: Surgical automation is being increasingly studied, yet bridging visual scene understanding with autonomous actio…
PreSIST: Vision-Language-Informed Object Persistence Prediction in Open-World Scenes
arXiv:2607.04057v1 Announce Type: cross Abstract: Robots deployed over long periods must reason about environments that change over time. Existing long-term per…
A User-driven Design Framework for Robotaxi
arXiv:2602.19107v3 Announce Type: replace Abstract: Robotaxis are emerging as a promising form of urban mobility, but removing human drivers fundamentally resha…
$\mathcal{P}^3$: Toward Versatile Embodied Agents
arXiv:2508.07033v2 Announce Type: replace Abstract: Embodied agents have shown promising generalization capabilities across diverse physical environments, makin…
3D Cal: An Open-Source Software Library for Depth Reconstruction on Vision-Based Tactile Sensors
arXiv:2511.03078v3 Announce Type: replace Abstract: Tactile sensing plays a key role in enabling dexterous and reliable robotic manipulation, but realizing this…
TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training
arXiv:2607.02840v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown promising generalization in robotic manipulation, but they still …
MorphQuad: Morphable Quadrotor for Superhuman Maneuverability, Manipulation, and Resiliency
arXiv:2607.02764v1 Announce Type: new Abstract: Infrastructure maintenance, contact-based inspection, and emergency response can benefit from aerial vehicles th…
Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis
arXiv:2607.05348v1 Announce Type: cross Abstract: Open-vocabulary 3D scene understanding aims to segment 3D scenes beyond predefined categories by transferring …
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
arXiv:2607.02642v1 Announce Type: new Abstract: Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficien…
SLAM: Structured and Localized Analytic Manifold Adaptation for Lifelong VPR
arXiv:2607.04764v1 Announce Type: new Abstract: Visual Place Recognition (VPR) in lifelong deployment requires continuous adaptation to new environments without…
Learning 3D Affordances for Blade Insertion in Cluttered Stowing
arXiv:2607.02549v1 Announce Type: cross Abstract: Many manipulation tasks require reasoning about free-space affordances: discovering volumes where an extended …
WSA$_1$: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control
arXiv:2607.03941v1 Announce Type: new Abstract: Recent advances in embodied AI have established robot foundation models (RFMs) as the dominant approach for gene…
FLOAT Drone for Physical Interaction: Lateral Airflow Reduction, Wrench Modeling, and Adaptive Control
arXiv:2607.04260v1 Announce Type: new Abstract: Aerial physical interaction represents a promising direction for next-generation unmanned aerial vehicles (UAVs)…
Real-World Perturbation Testing of Autonomous Driving Systems
arXiv:2607.04953v1 Announce Type: cross Abstract: Autonomous Driving Systems (ADS) must operate reliably under diverse conditions, yet representative data for r…
Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue
arXiv:2509.15061v3 Announce Type: replace Abstract: Embodied agents are intelligent systems designed to perceive, reason, and act within the physical world. Whi…
KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation
arXiv:2607.04652v1 Announce Type: new Abstract: Learning manipulation from few demonstrations requires visual priors that capture not only where to interact, bu…
Learning to Visually Connect Actions and their Effects
arXiv:2401.10805v4 Announce Type: replace-cross Abstract: We introduce the novel concept of visually Connecting Actions and Their Effects (CATE) in video unders…
ObjRetarget: An Object-Aware Motion Retargeting Framework with Anthropomorphic Arm Constraints and Polyhedral Hand Modeling
arXiv:2607.03828v1 Announce Type: new Abstract: Learning robot dexterous manipulation from human manipulation videos requires reliably retargeting human intent …
Verifier-free Test-Time Sampling for Vision-Language-Action Models
arXiv:2510.05681v2 Announce Type: replace Abstract: Vision-Language-Action models (VLAs) have demonstrated remarkable performance in robot control. However, the…