Who reported this story?

This story was reported by arXiv cs.RO.

Robotics

In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics

Robos News Newsroom

Editorial Desk

2026-06-26 · 2 min read

Published June 26, 2026 · Category: Robotics

Overview

arXiv:2606.26981v1 Announce Type: new Abstract: Synthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent trade-off between semantic fidelity and physical realism. Large language model (LLM)-based approaches can interpret diverse open-vocabulary instructions and compose high-level action plans, but they often generate motions that violate physical constraints. Physics-aware models improve realism through simulation or control, but they struggle with semantic complexity, fine-grained instructions, and novel concepts. To address this gap, we propose In-Context Model Predictive Generation (ICMPG), a framework that integrates language-model planning with inference-time physical feedback. ICMPG reformulates motion synthesis as a Model Predictive Control (MPC)-like process with two modules. The Context-Aware Motion Generation (CAMG) module uses an LLM as a planner to decompose textual commands and generate candidate motion sequences from motion tokens. The Model Predictive Generation (MPG) module evaluates these candidates through physical simulation and semantic alignment, estimates a composite reward, and selects the best sequence to guide subsequent generation steps. Unlike open-loop generation, this closed-loop refinement enables ICMPG to adapt motions to both the input semantics and the simulated physical environment without task-specific policy retraining. Extensive experiments across standard and zero-shot open-vocabulary settings show that ICMPG generalizes robustly to diverse commands and produces motions that are more physically plausible and semantically faithful than representative baselines on the evaluated benchmarks. The framework bridges semantic interpretation and physical simulation while remaining flexible enough to incorporate different LLM backbones, enabling more versatile and controllable text-driven motion synthesis.

Source

Originally published at arxiv.org.

Source: https://arxiv.org/abs/2606.26981

Robos News Newsroom

Robos News covers markets, crypto and commodities for Asia & the Middle East — tier-1 desk research, AI-driven analysis, institutional-grade data. Tip our newsroom: [email protected]

Email the newsroom →

Disclaimer: This article is for informational purposes only and does not constitute investment advice. Data may be delayed up to 15 minutes. Past performance is not indicative of future results. Consult a licensed financial advisor before making investment decisions.

In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics

Overview

Source

Related Articles

Related Stories

Overview

Source

Related Articles

Related Stories

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?

Monte Carlo Tree Search with Tensor Factorization for Optimization Problems in Robotics

A System for Fast, Resilient, and Adaptable Loco-Manipulation Behaviors on Humanoid Robots

FC-Vision: Real-Time Visibility-Aware Replanning for Occlusion-Free Aerial Target Structure Scanning in Unknown Environments

Cookie Preferences