Learning to Stack: Cube-Stacking Imitation Learning from Virtual Reality Demonstrations
arXiv:2609.19040v1 Announce Type: new Abstract: Imitation learning is attractive for robot manipulation, but collecting demonstrations remains a bottleneck for multi-stage tasks requiring repeated scene resets. This work presents a virtual-reality data-collection pipeline for cube-stacking with a custom 5-DoF arm in NVIDIA Isaac Sim and Isaac Lab. Using an HTC Vive Pro 2, Manus Quantum gloves, and OpenXR, an operator provides SE(3) end-effector commands to generate task demonstrations. The prop
Overview
arXiv:2609.19040v1 Announce Type: new Abstract: Imitation learning is attractive for robot manipulation, but collecting demonstrations remains a bottleneck for multi-stage tasks requiring repeated scene resets. This work presents a virtual-reality data-collection pipeline for cube-stacking with a custom 5-DoF arm in NVIDIA Isaac Sim and Isaac Lab. Using an HTC Vive Pro 2, Manus Quantum gloves, and OpenXR, an operator provides SE(3) end-effector commands to generate task demonstrations. The proposed framework separates demonstration collection from dataset construction by replaying recorded trajectories, converting task-space commands into joint-space actions, and re-rendering demonstrations with updated sensor or state configurations. This allows previously collected demonstrations to be reused for new observation and action spaces without repeating human teleoperation. The task requires stacking the red cube on the blue cube and the green cube on the red cube, with randomized cube placement. In 30 minutes, 200 virtual demonstrations were collected, compared with 45 real-world demonstrations, and Isaac Mimic generated 100 additional samples. A behavior-cloning policy was trained from the virtual demonstrations using LeRobot-style dual-camera observations and evaluated in simulation.
Source
Originally published at arxiv.org.
Related Articles
Source: https://arxiv.org/abs/2609.19040