When Does Test-Time Physical Diagnosis Pay? A Frozen Policy Buys Evidence It Never Reads
arXiv:2609.22299v1 Announce Type: new Abstract: When a robot faces unfamiliar physical conditions, a common approach is to collect evidence about what changed and adapt. For such diagnosis to improve behavior, six ordered empirical conditions must hold: a meaningful reference, identifiability of the physical condition, use of the acquired evidence, decision value, selection value over a fixed alternative, and safe realization. We test this chain in controlled and public environments. It holds e
Overview
arXiv:2609.22299v1 Announce Type: new Abstract: When a robot faces unfamiliar physical conditions, a common approach is to collect evidence about what changed and adapt. For such diagnosis to improve behavior, six ordered empirical conditions must hold: a meaningful reference, identifiability of the physical condition, use of the acquired evidence, decision value, selection value over a fixed alternative, and safe realization. We test this chain in controlled and public environments. It holds end to end in our controlled environments. After transfer to unseen mechanisms, however, it breaks at evidence use. On decisions requiring the full trace, the frozen decoder does not change its choice. A linear model using only trace increments recovers the correct choice on mechanisms excluded from fitting, showing that the trace is informative but unused. The failure is concentrated at the richest evidence level: those decisions fall to chance, while decisions settled with lower-cost evidence remain correct, a split hidden by aggregate accuracy. The same chain can fail at other links in public environments. Successful physical identification therefore guarantees neither evidence use nor useful adaptation; evaluation should identify where the chain breaks rather than rely on recovery accuracy or aggregate performance alone.
Source
Originally published at arxiv.org.
Related Articles
Source: https://arxiv.org/abs/2609.22299