ImagiNav: Scalable Embodied Navigation via Generative Visual Prediction and Inverse Dynamics
arXiv:2603.13833v2 Announce Type: replace Abstract: Enabling robots to navigate open-world environments via natural language is critical for general-purpose autonomy. Yet, Vision-Language Navigation has relied on end-to-end policies trained on expensive, embodiment-specific robot data. While recent foundation models trained on vast simulation data show promise, the challenge of scaling and generalizing persists due to the limited scene diversity and visual fidelity in simulation. To address thi
Overview
arXiv:2603.13833v2 Announce Type: replace Abstract: Enabling robots to navigate open-world environments via natural language is critical for general-purpose autonomy. Yet, Vision-Language Navigation has relied on end-to-end policies trained on expensive, embodiment-specific robot data. While recent foundation models trained on vast simulation data show promise, the challenge of scaling and generalizing persists due to the limited scene diversity and visual fidelity in simulation. To address this gap, we propose ImagiNav, a novel hierarchical paradigm that formulates navigation in visual space. Instead of predicting waypoints, ImagiNav synthesizes a future egocentric video conditioned on language instructions, serving as a high-level plan, interpreted by an inverse dynamics model to extract metric trajectories for execution. By decoupling planning from robot actuation, the paradigm enables direct utilization of diverse in-the-wild navigation videos. To support this, we develop an auto-labeling data pipeline that enhances motion annotation accuracy. ImagiNav demonstrates strong zero-shot transfer to robot navigation without requiring robot demonstrations, paving the way for generalist robots that learn navigation directly from unlabeled, open-world data. The project page is available at: https://j1dan.github.io/ImagiNav
Source
Originally published at arxiv.org.
Related Articles
Source: https://arxiv.org/abs/2603.13833



