One Word, Different Action: A Real-Robot Benchmark for Language-Conditioned Embodied Reasoning
arXiv:2609.05260v1 Announce Type: new Abstract: Natural-language instruction changes can directly alter robot behavior. A reliable embodied system should preserve its action when the task is unchanged and update it correctly when the task itself changes. We introduce One Word, Different Action, a real-robot benchmark built on physical decision states and executable actions, using task-preserving and task-changing instruction pairs to jointly evaluate Decision Invariance and Decision Sensitivity
Overview
arXiv:2609.05260v1 Announce Type: new Abstract: Natural-language instruction changes can directly alter robot behavior. A reliable embodied system should preserve its action when the task is unchanged and update it correctly when the task itself changes. We introduce One Word, Different Action, a real-robot benchmark built on physical decision states and executable actions, using task-preserving and task-changing instruction pairs to jointly evaluate Decision Invariance and Decision Sensitivity, with further evaluation under multi-constraint reasoning and real-RGB grounding. Experiments show that modern models are near saturation on single-constraint instruction changes, yet several models degrade noticeably when multiple task constraints must be integrated into one executable decision. These results suggest that the more salient remaining challenge is no longer recognizing an isolated instruction change, but reliably composing multiple task requirements into a correct robot action decision.
Source
Originally published at arxiv.org.
Related Articles
- FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models
- CoLMIN: LLM-based Multi-Decision Path Negotiation for Cooperative Autonomous Driving
- Strict Modes Everywhere - Bringing Order Into Dynamics of Mechanical Systems by a Potential Compatible With the Geodesic Flow
Source: https://arxiv.org/abs/2609.05260


