ReMemNav: Memory-Based Decision Correction and Target Verification for Zero-Shot Object Navigation
arXiv:2603.26788v3 Announce Type: replace Abstract: Zero-shot object navigation requires agents to locate unseen targets in unfamiliar environments without prior maps or task-specific training. Despite the commonsense reasoning ability of vision-language models (VLMs), existing mapless navigators often suffer from repeated exploration and premature stopping caused by limited historical context and false-positive target predictions. We propose ReMemNav, a training-free framework that combines li
Overview
arXiv:2603.26788v3 Announce Type: replace Abstract: Zero-shot object navigation requires agents to locate unseen targets in unfamiliar environments without prior maps or task-specific training. Despite the commonsense reasoning ability of vision-language models (VLMs), existing mapless navigators often suffer from repeated exploration and premature stopping caused by limited historical context and false-positive target predictions. We propose ReMemNav, a training-free framework that combines lightweight semantic grounding, memory-based decision correction, and target verification. RAM-derived semantic priors guide panoramic direction selection, while a geometry-triggered correction mechanism retrieves historical scene descriptions when the selected direction points toward previously visited regions. When a target is predicted to be visible, ReMemNav verifies the corresponding single-view observation before the final approach. Depth-based action sampling then converts the high-level decision into collision-free motion. Experiments on HM3D and MP3D show that ReMemNav achieves higher success rates and path efficiency than existing training-free zero-shot baselines. Specifically, ReMemNav achieves absolute SR/SPL gains of 1.7%/5.3% on HM3D v0.1, 18.2%/11.1% on HM3D v0.2, and 8.7%/6.9% on MP3D.
Source
Originally published at arxiv.org.
Related Articles
Source: https://arxiv.org/abs/2603.26788