Papers
Topics
Authors
Recent
Search
2000 character limit reached

TEXEDO: Controller-Aware Humanoid Motion

Updated 6 July 2026
  • TEXEDO is a controller-aware test-time scaling framework that transforms a frozen text-conditioned generator into a deployment-ready system via candidate sampling and dynamic feasibility verification.
  • It employs dual verifiers—one enforcing a hard dynamic feasibility gate and another ensuring semantic alignment—to select motions that are both safe and faithful to the input text.
  • The best-of-N selection process, validated on Unitree G1, significantly improves success rates and motion quality, making it a practical approach for language-conditioned humanoid robotics.

Searching arXiv for TEXEDO and closely related humanoid motion generation papers. Search query: TEXEDO controller-aware language-conditioned humanoid motion generation arXiv TEXEDO is a controller-aware test-time scaling framework for language-conditioned humanoid motion generation. It upgrades a frozen text-conditioned generator into a deployment-oriented system by sampling multiple candidate motions for a prompt and selecting the candidate that is both physically executable by a specific whole-body tracking controller and semantically aligned with the text. In the reported instantiation, TEXEDO operates on Unitree G1 whole-body trajectories, uses a dynamic feasibility verifier distilled from controller rollouts together with a semantic alignment verifier in a learned text–motion co-embedding space, and enforces a hard feasibility gate before semantic re-ranking (Cao et al., 22 Jun 2026).

1. Problem setting and design rationale

Language-conditioned whole-body motion generation is motivated by the prospect of using text as an interface for programming humanoid robots. The central difficulty addressed by TEXEDO is that most generators are trained on human motion corpora retargeted to robot morphologies, whereas deployment is governed by a particular whole-body tracking controller whose executable envelope depends on balance, contact timing, actuation limits, and controller-specific failure modes. A generated motion can therefore be semantically plausible while remaining difficult or impossible for the robot to execute (Cao et al., 22 Jun 2026).

The framework is organized as “Text-See-Do.” In the “TEXT” stage, a pretrained text-conditioned generator GG maps a prompt tt to a set of sampled candidate motions

MN(t)={mi}i=1N,M_N(t)=\{m_i\}_{i=1}^N,

where m∈RT×Dm \in \mathbb{R}^{T\times D}. For Unitree G1, the state dimension is D=36D=36, comprising root position, root quaternion, and 29 joints at $50$ Hz. In the “SEE” stage, each candidate is scored by a dynamic feasibility verifier Rdyn(m)R_{\text{dyn}}(m) and a semantic alignment verifier Rtext(t,m)R_{\text{text}}(t,m). In the “DO” stage, dynamic feasibility is treated as a hard constraint and semantic alignment as the objective: max⁡m∈MN(t)ssem(m,t)subject tofdyn(m)≥τ.\max_{m \in M_N(t)} s_{\text{sem}}(m,t)\quad \text{subject to}\quad f_{\text{dyn}}(m)\ge \tau. The deployed form is

R(m,t)={ssem(m,t),if fdyn(m)≥τ, −∞,otherwise.R(m,t)= \begin{cases} s_{\text{sem}}(m,t), & \text{if } f_{\text{dyn}}(m)\ge \tau,\ -\infty, & \text{otherwise.} \end{cases}

The reported threshold is tt0, chosen on a validation split to minimize false positives, and if no candidate passes the gate the method falls back to tt1. This fallback is reported to occur for only tt2 of prompts (Cao et al., 22 Jun 2026).

This design reframes motion generation as a selection problem rather than a retraining problem. A plausible implication is that TEXEDO is especially useful when the generator and the controller are already available but mismatched, because the method leaves both frozen and inserts grounded verification at inference time.

2. Controller-aware test-time scaling pipeline

TEXEDO’s defining operation is best-of-tt3 selection under grounded constraints. As tt4 increases, the system exploits the diversity of the generator’s output distribution rather than attempting to alter that distribution. The paper reports a monotonic test-time scaling law: using the appropriate verifier, larger candidate pools improve executability and semantic quality without requiring a stronger generator (Cao et al., 22 Jun 2026).

The practical selection algorithm is straightforward. For a prompt tt5, the system samples tt6 independent motions, decodes them to continuous trajectories, computes tt7 and tt8, filters candidates with tt9, and selects the feasible candidate with maximal MN(t)={mi}i=1N,M_N(t)=\{m_i\}_{i=1}^N,0. If the feasible set is empty, it returns the dynamically safest motion according to MN(t)={mi}i=1N,M_N(t)=\{m_i\}_{i=1}^N,1. This produces a conservative failure mode: semantics are sacrificed only when no candidate satisfies the feasibility gate.

The runtime cost is reported for an A100 40GB GPU with batched sampling and scoring. Sampling MN(t)={mi}i=1N,M_N(t)=\{m_i\}_{i=1}^N,2 candidates takes MN(t)={mi}i=1N,M_N(t)=\{m_i\}_{i=1}^N,3 s, scoring MN(t)={mi}i=1N,M_N(t)=\{m_i\}_{i=1}^N,4 candidates takes approximately MN(t)={mi}i=1N,M_N(t)=\{m_i\}_{i=1}^N,5 ms for MN(t)={mi}i=1N,M_N(t)=\{m_i\}_{i=1}^N,6 and MN(t)={mi}i=1N,M_N(t)=\{m_i\}_{i=1}^N,7 ms for MN(t)={mi}i=1N,M_N(t)=\{m_i\}_{i=1}^N,8, and the end-to-end latency for MN(t)={mi}i=1N,M_N(t)=\{m_i\}_{i=1}^N,9 is approximately m∈RT×Dm \in \mathbb{R}^{T\times D}0 s per prompt. These numbers define the practical operating point used throughout the main experiments (Cao et al., 22 Jun 2026).

A central conceptual feature is the asymmetry between the two verifiers. Dynamic feasibility is not used as a soft preference; it is a safety gate. Semantic alignment is not used to rescue unsafe trajectories; it only ranks candidates inside the feasible set. This sharply separates “can the robot execute it?” from “does it match the prompt?”

3. Dynamic feasibility verifier

The dynamic feasibility verifier is distilled from whole-body tracking rollouts. Reference motions are executed by a frozen SONIC whole-body tracker in MuJoCo until clip completion or early termination. Each rollout yields three supervisory signals: success m∈RT×Dm \in \mathbb{R}^{T\times D}1, progress ratio m∈RT×Dm \in \mathbb{R}^{T\times D}2, and normalized tracking quality m∈RT×Dm \in \mathbb{R}^{T\times D}3. The quality term is

m∈RT×Dm \in \mathbb{R}^{T\times D}4

where the velocity and acceleration errors are normalized by 95th-percentile constants (Cao et al., 22 Jun 2026).

These rollout signals are collapsed into an oracle quality

m∈RT×Dm \in \mathbb{R}^{T\times D}5

with m∈RT×Dm \in \mathbb{R}^{T\times D}6 and m∈RT×Dm \in \mathbb{R}^{T\times D}7. The constraint

m∈RT×Dm \in \mathbb{R}^{T\times D}8

guarantees success-first ordering: every successful rollout outranks every failed rollout, while failures remain graded by progress and quality. TEXEDO’s predicted dynamic score is

m∈RT×Dm \in \mathbb{R}^{T\times D}9

The verifier input at each frame is a 94-dimensional feature vector: 36-dimensional raw pose together with first and second finite differences, specifically root dynamics 7, joint position 29, joint velocity 29, and joint acceleration 29. The temporal encoder is a 4-layer causal Transformer with D=36D=360, 4 heads, pre-LayerNorm, and mean-attention pooling, followed by three heads predicting D=36D=361, D=36D=362, and D=36D=363. Training uses

D=36D=364

with D=36D=365 and D=36D=366 (Cao et al., 22 Jun 2026).

Executability is grounded in controller-aware termination criteria: anchor height deviation D=36D=367 with D=36D=368 m, anchor tilt exceeding D=36D=369, and end-effector clearance exceeding $50$0 m. These conditions encode balance and contact constraints, while $50$1 captures bandwidth limits through velocity and acceleration tracking. On held-out data, the success head reaches AUROC $50$2 and AUPRC $50$3, with failure recall $50$4. Rank agreement with the simulator oracle is reported as within-prompt Kendall $50$5 overall and $50$6 on mixed success/failure prompts (Cao et al., 22 Jun 2026).

4. Semantic alignment verifier and motion generator

The semantic verifier is a learned text–motion co-embedding model trained directly on robot-skeleton motions. The motion encoder consists of a MovementConvEncoder with two temporal Conv1d downsampling layers followed by a BiGRU, yielding a 512-dimensional motion embedding $50$7. Root $50$8 positions are replaced by frame-to-frame velocities to enforce translation invariance. The text encoder uses 300-dimensional word embeddings, 15-dimensional part-of-speech embeddings, and a BiGRU to produce a 512-dimensional text embedding $50$9 (Cao et al., 22 Jun 2026).

Training uses an all-pairs margin loss

Rdyn(m)R_{\text{dyn}}(m)0

with margin Rdyn(m)R_{\text{dyn}}(m)1. At test time,

Rdyn(m)R_{\text{dyn}}(m)2

A motion-to-text retrieval sanity check on 9,116 held-out pairs with 32 distractors yields Rdyn(m)R_{\text{dyn}}(m)3 and Rdyn(m)R_{\text{dyn}}(m)4, compared with a random baseline of approximately Rdyn(m)R_{\text{dyn}}(m)5 (Cao et al., 22 Jun 2026).

The default generator in the reported system is FSQ-GPT. Motions are represented at 50 Hz with 36-dimensional frames. An FSQ VAE with two stride-2 temporal downsamples compresses a sequence to Rdyn(m)R_{\text{dyn}}(m)6 tokens. The FSQ code levels are Rdyn(m)R_{\text{dyn}}(m)7, giving Rdyn(m)R_{\text{dyn}}(m)8 discrete codes. Reconstruction uses SmoothL1 losses on pose, velocity, and acceleration with weights Rdyn(m)R_{\text{dyn}}(m)9 for root position, Rtext(t,m)R_{\text{text}}(t,m)0 for root quaternion, Rtext(t,m)R_{\text{text}}(t,m)1 for joint position, Rtext(t,m)R_{\text{text}}(t,m)2 for root velocity, Rtext(t,m)R_{\text{text}}(t,m)3 for joint velocity, Rtext(t,m)R_{\text{text}}(t,m)4 for root acceleration, and Rtext(t,m)R_{\text{text}}(t,m)5 for joint acceleration. The LLM is Flan-T5-base, approximately Rtext(t,m)R_{\text{text}}(t,m)6M parameters, with vocabulary extended by the motion codes and special tokens, trained with teacher-forced cross-entropy (Cao et al., 22 Jun 2026).

Training data combine AMASS with HumanML3D captions, retargeted to Unitree G1, together with the CLAW corpus, using an 8:1:1 train/validation/test split and 9,116 held-out prompts. Sampling uses ancestral decoding with temperature Rtext(t,m)R_{\text{text}}(t,m)7, top-Rtext(t,m)R_{\text{text}}(t,m)8, and top-Rtext(t,m)R_{\text{text}}(t,m)9. Because the verifiers operate on decoded trajectories rather than generator internals, TEXEDO transfers in plug-and-play fashion to unseen generators such as Kimodo (Cao et al., 22 Jun 2026).

5. Evaluation, scaling behavior, and deployment results

Evaluation uses Unitree G1 with the frozen SONIC controller. Dynamics are measured with success rate, MPJPE in millimeters, max⁡m∈MN(t)ssem(m,t)subject tofdyn(m)≥τ.\max_{m \in M_N(t)} s_{\text{sem}}(m,t)\quad \text{subject to}\quad f_{\text{dyn}}(m)\ge \tau.0 in mm/framemax⁡m∈MN(t)ssem(m,t)subject tofdyn(m)≥τ.\max_{m \in M_N(t)} s_{\text{sem}}(m,t)\quad \text{subject to}\quad f_{\text{dyn}}(m)\ge \tau.1, max⁡m∈MN(t)ssem(m,t)subject tofdyn(m)≥τ.\max_{m \in M_N(t)} s_{\text{sem}}(m,t)\quad \text{subject to}\quad f_{\text{dyn}}(m)\ge \tau.2 in mm/frame, and max⁡m∈MN(t)ssem(m,t)subject tofdyn(m)≥τ.\max_{m \in M_N(t)} s_{\text{sem}}(m,t)\quad \text{subject to}\quad f_{\text{dyn}}(m)\ge \tau.3. Semantic quality is measured with a VLM-as-Judge ensemble using GPT-5.5 and Gemini-2.5, with two rubrics and two frame samplings. The main comparison is between Base (max⁡m∈MN(t)ssem(m,t)subject tofdyn(m)≥τ.\max_{m \in M_N(t)} s_{\text{sem}}(m,t)\quad \text{subject to}\quad f_{\text{dyn}}(m)\ge \tau.4), max⁡m∈MN(t)ssem(m,t)subject tofdyn(m)≥τ.\max_{m \in M_N(t)} s_{\text{sem}}(m,t)\quad \text{subject to}\quad f_{\text{dyn}}(m)\ge \tau.5-only selection, max⁡m∈MN(t)ssem(m,t)subject tofdyn(m)≥τ.\max_{m \in M_N(t)} s_{\text{sem}}(m,t)\quad \text{subject to}\quad f_{\text{dyn}}(m)\ge \tau.6-only selection, full TEXEDO, and an Oracle upper bound based on simulator rollouts (Cao et al., 22 Jun 2026).

At max⁡m∈MN(t)ssem(m,t)subject tofdyn(m)≥τ.\max_{m \in M_N(t)} s_{\text{sem}}(m,t)\quad \text{subject to}\quad f_{\text{dyn}}(m)\ge \tau.7 with FSQ-GPT, the Base system achieves VLM max⁡m∈MN(t)ssem(m,t)subject tofdyn(m)≥τ.\max_{m \in M_N(t)} s_{\text{sem}}(m,t)\quad \text{subject to}\quad f_{\text{dyn}}(m)\ge \tau.8, Succ max⁡m∈MN(t)ssem(m,t)subject tofdyn(m)≥τ.\max_{m \in M_N(t)} s_{\text{sem}}(m,t)\quad \text{subject to}\quad f_{\text{dyn}}(m)\ge \tau.9, MPJPE R(m,t)={ssem(m,t),if fdyn(m)≥τ, −∞,otherwise.R(m,t)= \begin{cases} s_{\text{sem}}(m,t), & \text{if } f_{\text{dyn}}(m)\ge \tau,\ -\infty, & \text{otherwise.} \end{cases}0, R(m,t)={ssem(m,t),if fdyn(m)≥τ, −∞,otherwise.R(m,t)= \begin{cases} s_{\text{sem}}(m,t), & \text{if } f_{\text{dyn}}(m)\ge \tau,\ -\infty, & \text{otherwise.} \end{cases}1, R(m,t)={ssem(m,t),if fdyn(m)≥τ, −∞,otherwise.R(m,t)= \begin{cases} s_{\text{sem}}(m,t), & \text{if } f_{\text{dyn}}(m)\ge \tau,\ -\infty, & \text{otherwise.} \end{cases}2, and R(m,t)={ssem(m,t),if fdyn(m)≥τ, −∞,otherwise.R(m,t)= \begin{cases} s_{\text{sem}}(m,t), & \text{if } f_{\text{dyn}}(m)\ge \tau,\ -\infty, & \text{otherwise.} \end{cases}3. R(m,t)={ssem(m,t),if fdyn(m)≥τ, −∞,otherwise.R(m,t)= \begin{cases} s_{\text{sem}}(m,t), & \text{if } f_{\text{dyn}}(m)\ge \tau,\ -\infty, & \text{otherwise.} \end{cases}4-only improves executability to Succ R(m,t)={ssem(m,t),if fdyn(m)≥τ, −∞,otherwise.R(m,t)= \begin{cases} s_{\text{sem}}(m,t), & \text{if } f_{\text{dyn}}(m)\ge \tau,\ -\infty, & \text{otherwise.} \end{cases}5, MPJPE R(m,t)={ssem(m,t),if fdyn(m)≥τ, −∞,otherwise.R(m,t)= \begin{cases} s_{\text{sem}}(m,t), & \text{if } f_{\text{dyn}}(m)\ge \tau,\ -\infty, & \text{otherwise.} \end{cases}6, R(m,t)={ssem(m,t),if fdyn(m)≥τ, −∞,otherwise.R(m,t)= \begin{cases} s_{\text{sem}}(m,t), & \text{if } f_{\text{dyn}}(m)\ge \tau,\ -\infty, & \text{otherwise.} \end{cases}7, R(m,t)={ssem(m,t),if fdyn(m)≥τ, −∞,otherwise.R(m,t)= \begin{cases} s_{\text{sem}}(m,t), & \text{if } f_{\text{dyn}}(m)\ge \tau,\ -\infty, & \text{otherwise.} \end{cases}8, and R(m,t)={ssem(m,t),if fdyn(m)≥τ, −∞,otherwise.R(m,t)= \begin{cases} s_{\text{sem}}(m,t), & \text{if } f_{\text{dyn}}(m)\ge \tau,\ -\infty, & \text{otherwise.} \end{cases}9, but semantics drop to VLM tt00. tt01-only improves semantics to VLM tt02, with more modest dynamics gains: Succ tt03, MPJPE tt04, tt05, tt06, tt07. Full TEXEDO balances the two objectives: VLM tt08, Succ tt09, MPJPE tt10, tt11, tt12, and tt13 (Cao et al., 22 Jun 2026).

Relative to Base, TEXEDO yields tt14 Succ, tt15 mm MPJPE, tt16 mm/framett17 tt18, tt19 mm/frame tt20, tt21 in tt22, and tt23 VLM. The reported interpretation is that TEXEDO preserves most of the semantic benefit of tt24-only while retaining most of the executability benefit of tt25-only.

The same pattern persists under transfer and distribution shift. For unseen Kimodo, Base at tt26 obtains VLM tt27, Succ tt28, MPJPE tt29, tt30, tt31, tt32; with full TEXEDO at tt33, these become VLM tt34, Succ tt35, MPJPE tt36, tt37, tt38, and tt39. On BONES-SEED out-of-distribution prompts, Base reports VLM tt40, Succ tt41, MPJPE tt42, tt43, tt44, tt45, while TEXEDO reaches VLM tt46, Succ tt47, MPJPE tt48, tt49, tt50, and tt51 (Cao et al., 22 Jun 2026).

Real-world deployment on a Unitree G1 reports 30/30 successful executions for 30 prompts at tt52, with representative tracking averages MPJPE tt53 mm, tt54 mm/frame, and tt55 mm/framett56. The reported qualitative coverage includes locomotion, gestures, and compositional prompt combinations (Cao et al., 22 Jun 2026).

6. Limitations, scope, and research significance

TEXEDO introduces a specific form of controller awareness, but its guarantees remain conditional on the learned verifier and the downstream tracker. The paper identifies a latency–quality trade-off controlled by tt57, verifier dependence on the controller used to label rollouts, the possibility of rare but consequential false positives, weaker semantic generalization on out-of-distribution prompts, and the open-loop character of the current selection procedure. The dynamic verifier is controller-specific; swapping controllers requires relabeling via new rollouts, although the underlying motion corpora can be reused. Multi-robot portability is therefore not automatic (Cao et al., 22 Jun 2026).

The framework is also informative as a methodological statement. It does not align generator and tracker by joint retraining, nor does it fine-tune the controller to absorb generator errors. Instead, it keeps both components frozen and inserts a controller-grounded verifier at runtime. This places TEXEDO within a broader class of test-time scaling methods, but with an unusual reward structure in which safety-like feasibility is hard-gated and semantics are optimized only after feasibility has been established.

Within language-conditioned motion generation, this architecture addresses a common failure mode of kinematically plausible but dynamically untrackable outputs. The reported avoided errors include foot scuffing that leads to trips, excessive pelvis pitch or roll, overly rapid limb swings that saturate actuators, and contact timing mismatches. This suggests that TEXEDO’s main contribution is not a new generative prior, but an executable selection policy grounded in the dynamics of a particular robot–controller pair (Cao et al., 22 Jun 2026).

The resulting picture is that TEXEDO converts extra inference-time samples into a single reference trajectory that is both executable and semantically faithful. In the reported experiments, that conversion improves simulation metrics, transfers across generators, generalizes to out-of-distribution prompts, and carries through to real-world humanoid deployment without modifying either the generator or the controller.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to TEXEDO.