Privileged observations enable rapid and reliable policy discovery directly in the physical world
arXiv:2512.08463v2 Announce Type: replace-cross Abstract: We study how privileged information about a physical system affects the discovery of high-performing policies when training a reinforcement learning agent directly in the physical world. We let the agent control a cylinder in a tabletop water channel to maximize or minimize drag. The flow is chaotic and difficult to model or simulate accurately and good strategies are not obvious beforehand. Decades-old experimental studies provide recip
Overview
arXiv:2512.08463v2 Announce Type: replace-cross Abstract: We study how privileged information about a physical system affects the discovery of high-performing policies when training a reinforcement learning agent directly in the physical world. We let the agent control a cylinder in a tabletop water channel to maximize or minimize drag. The flow is chaotic and difficult to model or simulate accurately and good strategies are not obvious beforehand. Decades-old experimental studies provide recipes for simple, high-performance, periodic open-loop policies that increase drag by 26.6% +/- 0.7% and reduce it by 29.7% +/- 1.3%. With dense observations of the cylinder wake, the agent learns within tens of minutes policies that increase drag by 25.5% +/- 0.9% and reduce it by 32.4% +/- 1.6%. We record action trajectories during online policy execution and replay them open loop as fixed sequences; these replays increase drag by 23.2% +/- 2.2% and reduce it by 32.1% +/- 3.0%. However, when we withhold flow observations during training, the agent still learns to decrease drag by 31.4% +/- 2.1%, but no run learns to increase it (1.8% +/- 4.3% drag increase). Our physical experiments demonstrate an extreme case: privileged observations can be decisive for policy discovery even when the resulting actions can be replayed in open loop.
Source
Originally published at arxiv.org.
Related Articles
Source: https://arxiv.org/abs/2512.08463