Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesScaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning
arXiv:2608.15863v1 Announce Type: new Abstract: Operating household appliances requires long-horizon planning that is state-dependent and robust to disturbances…
RAPAC-DP: Response-Aligned Pending-Action Compensation for Diffusion Policies under Delayed Execution
arXiv:2608.15924v1 Announce Type: new Abstract: Cloud-side inference gives imitation-learning policies access to greater computational resources, but communicat…
OccamView: Object-Conditioned View Selection for Frame-Budgeted Active 3D Gaussian Reconstruction
arXiv:2608.16499v1 Announce Type: new Abstract: Active 3D Gaussian reconstruction fundamentally relies on selecting informative next-best views under limited se…
ViHaTeleop: A Low-Cost, Lightweight Visual-Haptic Teleoperation System for Dexterous Manipulation Learning
arXiv:2608.16572v1 Announce Type: new Abstract: Learning from demonstration is a promising approach for dexterous manipulation, but collecting high-quality cont…
Orbit-Planner: Towards Latent World Models for On-Orbit Obstacle Avoidance of Satellite Agents
arXiv:2608.16651v1 Announce Type: new Abstract: Satellite agents for on-orbit navigation tasks need to predict collision risks using limited onboard observation…
H-PAC Hand: Control-Oriented Modeling and Tendon-Elasticity Compensation for an Underactuated Robotic Hand
arXiv:2608.16712v1 Announce Type: new Abstract: Underactuated tendon-driven hands offer compact actuation and passive compliance, but tendon elongation under re…
MatchingPolicy: Correspondence-Aware Policy Enables Cross-Object In-Context Learning
arXiv:2608.16715v1 Announce Type: new Abstract: In-context imitation learning enables few-shot policy generalization but struggles to maintain performance on un…
Adaptive Repulsive Pheromone Clustering for Foraging Robot Swarms
arXiv:2608.16822v1 Announce Type: new Abstract: The Central Place Foraging Algorithm (CPFA) combines site fidelity, pheromone-guided navigation, and uninformed …
FlexWorm: Primitive-augmented Hybrid Contact-motion Planning for Suction-based Multi-segment Deformable Robots
arXiv:2608.16853v1 Announce Type: new Abstract: Multi-segment suction-based soft robots are promising for inspection and maintenance in confined or fragile envi…
Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory
arXiv:2608.16889v1 Announce Type: new Abstract: Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-actio…
Beam-Wise Statistical Background Subtraction for Static Roadside LiDAR: A Cross-Sensor Benchmark Study
arXiv:2608.14868v1 Announce Type: cross Abstract: Background subtraction is a key preprocessing step for infrastructure-based LiDAR perception, enabling efficie…
Admissibility-Preserving Control for Strict-Feedback Nonlinear Systems with Asymmetric Actuator Constraints
arXiv:2608.15375v1 Announce Type: cross Abstract: This paper develops Admissibility-Preserving Control (APC), a realization-centered safety-critical control fra…
Pluralistic Human-Robot Interaction: Designing for Robot Interaction with Diverse Communities
arXiv:2608.16049v1 Announce Type: cross Abstract: Social robots are being developed for homes, schools, and other environments where they will interact with div…
HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents
arXiv:2608.16447v1 Announce Type: cross Abstract: Long-horizon embodied tasks require LLM agents to iteratively decompose high-level goals, revise plans in resp…
Seeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning
arXiv:2510.03885v4 Announce Type: replace Abstract: In this paper, we demonstrate that mobile manipulation policies utilizing a 3D latent map achieve stronger s…
PerFACT: Motion Policy with LLM-Powered Dataset Synthesis and Fusion Action-Chunking Transformers
arXiv:2512.03444v2 Announce Type: replace Abstract: Deep learning methods have significantly enhanced motion planning for robotic manipulators by leveraging pri…
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
arXiv:2512.22983v2 Announce Type: replace Abstract: Recent advances in vision, language, and multimodal learning have significantly accelerated progress in robo…
Hybrid System Planning using a Mixed-Integer ADMM Heuristic and Hybrid Zonotopes
arXiv:2602.17574v2 Announce Type: replace Abstract: Embedded optimization-based planning for hybrid systems is challenging due to the use of mixed-integer progr…
Grounding Robot Generalization in Training Data via Retrieval-Augmented VLMs
arXiv:2603.11426v3 Announce Type: replace Abstract: Recent work on robot manipulation has advanced policy generalization to novel scenarios. However, it is ofte…
Morphology-Conditioned World Model for Cross-Embodiment Quadrupedal Locomotion
arXiv:2604.08780v2 Announce Type: replace Abstract: World models promise a paradigm shift in robotics, where an agent learns the physics of its environment once…
SADP: Subgoal-Aware Diffusion Policy for Long-Horizon Manipulation Learned from Foundation Model Generated Demonstrations
arXiv:2605.16871v2 Announce Type: replace Abstract: Long-horizon robot manipulation requires policies to coordinate multiple intermediate subgoals and determine…
The functional and temporal roles of gaze evolve across the phases and constraints of multi-stage robot-mediated manipulation
arXiv:2606.21920v2 Announce Type: replace Abstract: Goal-directed eye movements are a fundamental component of visuomotor control, enabling humans to anticipate…
VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models
arXiv:2605.29605v2 Announce Type: replace Abstract: Task-success confidence estimation for Vision-Language-Action (VLA) models provides a crucial task-level sig…
Learning Versatile Humanoid Manipulation with Touch Dreaming
arXiv:2604.13015v3 Announce Type: replace Abstract: Humanoid robots promise general-purpose assistance, yet real-world humanoid loco-manipulation remains challe…
I-Perceive: A Foundation Model for Active Perception with Language Instructions
arXiv:2603.00600v2 Announce Type: replace Abstract: Active perception - the ability of a robot to proactively select viewpoints to acquire task-relevant informa…
MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery
arXiv:2602.12407v3 Announce Type: replace Abstract: Background: Robot-assisted minimally invasive surgery (RMIS) research increasingly relies on multimodal data…
On Minimum Aerial Photographs for Planar Region Coverage: Hardness and Approximation
arXiv:2512.18268v4 Announce Type: replace Abstract: Aerial photography with drones often requires covering a planar region with a limited number of images while…
Fully distributed and resilient source seeking for robot swarms
arXiv:2410.15921v3 Announce Type: replace Abstract: Existing source-seeking algorithms for robot swarms typically require either direct gradient measurements or…
EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints
arXiv:2608.15502v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising foundation for Embodied AI, but their high inf…
FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge
arXiv:2608.15410v1 Announce Type: cross Abstract: Reasoning segmentation enables vision-language models (VLMs) to translate mission-relevant language requests i…