Research
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Frontier robotics research: arXiv cs.RO papers, embodied AI, VLA models, manipulation, navigation and learning systems.
Latest in Research
4787 storiesAn AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop
arXiv:2608.07542v1 Announce Type: cross Abstract: Autonomous research loops driven by large language models can run machine-learning experiments at scale but te…
Contraction Analysis of Holomorphic Dynamical Systems via the Intrinsic Kobayashi Metric
arXiv:2608.07551v1 Announce Type: cross Abstract: This paper studies incremental stability of holomorphic dynamical systems through the infinitesimal Kobayashi …
Exact Contraction Rates via the Berkson--Porta Representation: A Sharp Threshold and Its Herglotz-Kernel Obstruction
arXiv:2608.07552v1 Announce Type: cross Abstract: Semigroups of holomorphic self-maps of the unit disc with an interior fixed point are, by the classical Berkso…
Impact of Dataset Composition on Embedded Real-Time UAV Wildfire Detection Using Compact YOLO Models
arXiv:2608.07554v1 Announce Type: cross Abstract: The development of vision-based wildfire detection systems for unmanned aerial vehicles is constrained by the …
The Field Knows: Cross-Dimensional Geometry from Navigation to Black Holes
arXiv:2608.07566v1 Announce Type: cross Abstract: We introduce a continuous metric field framework trained by a single causal contrastive loss. The framework en…
Multimodal Skin Lesion Classification with Swin Transformer and Clinical Metadata Fusion
arXiv:2608.07574v1 Announce Type: cross Abstract: Skin lesion classification plays an important role in supporting the early diagnosis of skin cancer. However, …
Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects
arXiv:2608.07577v1 Announce Type: cross Abstract: A closed-set detector for autonomous driving must assign every object one of a fixed set of labels. On an obje…
CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models
arXiv:2608.07621v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous dr…
LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation
arXiv:2608.07746v1 Announce Type: cross Abstract: Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level…
V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control
arXiv:2608.07870v1 Announce Type: cross Abstract: Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world …
GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning
arXiv:2608.07905v1 Announce Type: cross Abstract: Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to l…
Parameter-Dependent LMI Synthesis for Semi-Global Differential ISS Trajectory Tracking of Nonholonomic Mobile Robots Under Multiplicative Wheel Slip
arXiv:2608.08049v1 Announce Type: cross Abstract: This paper presents a parameter-dependent linear matrix inequality (LMI) framework for trajectory tracking of …
Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?
arXiv:2608.08077v1 Announce Type: cross Abstract: Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models …
Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-world maritime navigation scenarios
arXiv:2608.08281v1 Announce Type: cross Abstract: Recently, Large Language Models (LLMs) have shown considerable capability for situational understanding, reaso…
Ego-OSCAR: Egocentric Open source Stereo CAptuRe System
arXiv:2608.08285v1 Announce Type: cross Abstract: We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric d…
Machine-Learning-Based Diagnostic Framework for Passive Ultrasonic Detection of Railway Wheel Defects
arXiv:2608.08301v1 Announce Type: cross Abstract: Reliable identification of railway wheel defects is important for safety and maintenance. This study develops …
Preview-Based Relative-Motion Control of an Insertion Tool for Neural-Thread Placement in Pulsating Tissue
arXiv:2608.08860v1 Announce Type: cross Abstract: Robotic neural-thread placement requires regulating the insertion-tool tip relative to tissue that moves with …
Intuitive Hand Positional Guidance Using McKibben-Based Surface Tactile Sensations to Shoulder and Elbow
arXiv:2608.09167v1 Announce Type: cross Abstract: Hand positional guidance with intuitive perception is crucial for enhancing user interaction and task performa…
Intuitive Directional Sense Presentation to the Torso Using McKibben-Based Surface Haptic Sensation in Immersive Space
arXiv:2608.09177v1 Announce Type: cross Abstract: In recent years, systems that utilize immersive space have been developed in various fields. Immersive spaces …
Real-Time Nonlinear MPC via Sequential Quadratic Programming with Structure-Exploiting ADMM and Interior-Point Methods for Underactuated Double-Pendulum Swing-Up
arXiv:2608.09272v1 Announce Type: cross Abstract: The 4th "AI Olympics with RealAIGym" competition, to be held at IJCAI-ECAI 2026 in Bremen, challenges particip…
CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation
arXiv:2608.09296v1 Announce Type: cross Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, res…
Trajectory-Induced Self-Calibration for Hidden-Target Localization Through an Unknown-Pose Range-Bearing Relay
arXiv:2608.09464v1 Announce Type: cross Abstract: This paper studies hidden-target localization from range-bearing packets reported by a relay beacon whose glob…
A Height-Constrained 2-Point Minimal Solver for Pose Estimation from Active LED Markers with Event Cameras
arXiv:2608.09520v1 Announce Type: cross Abstract: In many autonomous applications requiring real-time localization, active marker-based systems are preferred du…
GenTrack3: Hybrid Stochastic-Deterministic Online Multi-Object Tracking with Cluster-Aware Association
arXiv:2608.09581v1 Announce Type: cross Abstract: Multi-object tracking (MOT) involves maintaining consistent target identities as objects dynamically enter and…
A Semantic Communication Approach to Fiducial Marker Processing in 5G-Enabled Edge SLAM
arXiv:2608.09620v1 Announce Type: cross Abstract: Autonomous robots increasingly rely on edge computing to offload computationally intensive perception tasks wh…
Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance
arXiv:2608.09628v1 Announce Type: cross Abstract: Collision avoidance systems are commonly used to avoid fragmentation events occurring in Low-Earth Orbit (LEO)…
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
arXiv:2503.22122v2 Announce Type: replace Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly fo…
X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation
arXiv:2505.11146v3 Announce Type: replace Abstract: Fine-grained facial expression transfer from humans to humanoid agents presents a unique pattern recognition…
Analysis and experiments of the dissipative Twistcar: direction reversal and asymptotic approximations
arXiv:2506.19112v3 Announce Type: replace Abstract: Underactuated wheeled vehicles are commonly studied as nonholonomic systems with periodic actuation. Twistca…
Calib3R: Hand-Eye Calibration and 3D Metric-Scaled Scene Reconstruction with 3D Foundation Models
arXiv:2509.08813v2 Announce Type: replace Abstract: Robots often rely on RGB images for tasks like manipulation. However, reliable interaction typically require…