---
title: 'TEXEDO: Controller-Aware Humanoid Motion'
url: https://www.emergentmind.com/topics/texedo
type: topic
---

# TEXEDO: Controller-Aware Humanoid Motion

Searching arXiv for TEXEDO and closely related humanoid motion generation papers.
Search query: TEXEDO controller-aware language-conditioned humanoid motion generation arXiv
TEXEDO is a controller-aware test-time scaling framework for language-conditioned humanoid motion generation. It upgrades a frozen text-conditioned generator into a deployment-oriented system by sampling multiple candidate motions for a prompt and selecting the candidate that is both physically executable by a specific whole-body tracking controller and semantically aligned with the text. In the reported instantiation, TEXEDO operates on Unitree G1 whole-body trajectories, uses a dynamic feasibility verifier distilled from controller rollouts together with a semantic alignment verifier in a learned text–motion co-embedding space, and enforces a hard feasibility gate before semantic re-ranking [2606.22998].

## 1. Problem setting and design rationale

Language-conditioned whole-body motion generation is motivated by the prospect of using text as an interface for programming humanoid robots. The central difficulty addressed by TEXEDO is that most generators are trained on human motion corpora retargeted to robot morphologies, whereas deployment is governed by a particular whole-body tracking controller whose executable envelope depends on balance, contact timing, actuation limits, and controller-specific failure modes. A generated motion can therefore be semantically plausible while remaining difficult or impossible for the robot to execute [2606.22998].

The framework is organized as “Text-See-Do.” In the “TEXT” stage, a pretrained text-conditioned generator \(G\) maps a prompt \(t\) to a set of sampled candidate motions
\[
M_N(t)=\{m_i\}_{i=1}^N,
\]
where \(m \in \mathbb{R}^{T\times D}\). For Unitree G1, the state dimension is \(D=36\), comprising root position, root quaternion, and 29 joints at \(50\) Hz. In the “SEE” stage, each candidate is scored by a dynamic feasibility verifier \(R_{\text{dyn}}(m)\) and a semantic alignment verifier \(R_{\text{text}}(t,m)\). In the “DO” stage, dynamic feasibility is treated as a hard constraint and semantic alignment as the objective:
\[
\max_{m \in M_N(t)} s_{\text{sem}}(m,t)\quad \text{subject to}\quad f_{\text{dyn}}(m)\ge \tau.
\]
The deployed form is
\[
R(m,t)=
\begin{cases}
s_{\text{sem}}(m,t), & \text{if } f_{\text{dyn}}(m)\ge \tau,\\
-\infty, & \text{otherwise.}
\end{cases}
\]
The reported threshold is \(\tau=0.8\), chosen on a validation split to minimize false positives, and if no candidate passes the gate the method falls back to \(\arg\max_i R_{\text{dyn}}(m_i)\). This fallback is reported to occur for only \(1.18\%\) of prompts [2606.22998].

This design reframes motion generation as a selection problem rather than a retraining problem. A plausible implication is that TEXEDO is especially useful when the generator and the controller are already available but mismatched, because the method leaves both frozen and inserts grounded verification at inference time.

## 2. Controller-aware test-time scaling pipeline

TEXEDO’s defining operation is best-of-\(N\) selection under grounded constraints. As \(N\) increases, the system exploits the diversity of the generator’s output distribution rather than attempting to alter that distribution. The paper reports a monotonic test-time scaling law: using the appropriate verifier, larger candidate pools improve executability and semantic quality without requiring a stronger generator [2606.22998].

The practical selection algorithm is straightforward. For a prompt \(t\), the system samples \(N\) independent motions, decodes them to continuous trajectories, computes \(R_{\text{dyn}}\) and \(R_{\text{text}}\), filters candidates with \(R_{\text{dyn}}<\tau\), and selects the feasible candidate with maximal \(R_{\text{text}}\). If the feasible set is empty, it returns the dynamically safest motion according to \(R_{\text{dyn}}\). This produces a conservative failure mode: semantics are sacrificed only when no candidate satisfies the feasibility gate.

The runtime cost is reported for an A100 40GB GPU with batched sampling and scoring. Sampling \(32\) candidates takes \(2.62\) s, scoring \(32\) candidates takes approximately \(15.12\) ms for \(R_{\text{dyn}}\) and \(11.93\) ms for \(R_{\text{text}}\), and the end-to-end latency for \(N=32\) is approximately \(3.5\) s per prompt. These numbers define the practical operating point used throughout the main experiments [2606.22998].

A central conceptual feature is the asymmetry between the two verifiers. Dynamic feasibility is not used as a soft preference; it is a safety gate. Semantic alignment is not used to rescue unsafe trajectories; it only ranks candidates inside the feasible set. This sharply separates “can the robot execute it?” from “does it match the prompt?”

## 3. Dynamic feasibility verifier

The dynamic feasibility verifier is distilled from whole-body tracking rollouts. Reference motions are executed by a frozen SONIC whole-body tracker in MuJoCo until clip completion or early termination. Each rollout yields three supervisory signals: success \(y_s \in \{0,1\}\), progress ratio \(q_g \in [0,1]\), and normalized tracking quality \(q_d \in [0,1]\). The quality term is
\[
q_d(m)=\tfrac{1}{2}\cdot \mathrm{clip}(1-\hat e_{acc},0,1)+\tfrac{1}{2}\cdot \mathrm{clip}(1-\hat e_{vel},0,1),
\]
where the velocity and acceleration errors are normalized by 95th-percentile constants [2606.22998].

These rollout signals are collapsed into an oracle quality
\[
Q^*(y_s,q_d,q_g)=y_s\cdot(1+\alpha q_d)+(1-y_s)\cdot \beta q_g q_d,
\]
with \(\alpha=0.4\) and \(\beta=0.6\). The constraint
\[
\beta < \frac{1}{1+\alpha}
\]
guarantees success-first ordering: every successful rollout outranks every failed rollout, while failures remain graded by progress and quality. TEXEDO’s predicted dynamic score is
\[
R_{\text{dyn}}(m)=Q^*(\hat y_s(m),\hat q_d(m),\hat q_g(m)).
\]

The verifier input at each frame is a 94-dimensional feature vector: 36-dimensional raw pose together with first and second finite differences, specifically root dynamics 7, joint position 29, joint velocity 29, and joint acceleration 29. The temporal encoder is a 4-layer causal Transformer with \(d_{\text{model}}=256\), 4 heads, pre-LayerNorm, and mean-attention pooling, followed by three heads predicting \(\hat y_s\), \(\hat q_d\), and \(\hat q_g\). Training uses
\[
\mathcal{L}=\mathrm{BCE}_{w+}(\hat y_s,y_s)+\lambda_d(\hat q_d-q_d)^2+\lambda_g\,\mathbb{1}[y_s=0]\,(\hat q_g-q_g)^2,
\]
with \(\lambda_d=0.6\) and \(\lambda_g=0.8\) [2606.22998].

Executability is grounded in controller-aware termination criteria: anchor height deviation \(|[p_{\text{ref},0}]_z-[p_{\text{rob},0}]_z|>h_a\) with \(h_a=0.25\) m, anchor tilt exceeding \(c_z=0.8\), and end-effector clearance exceeding \(h_e=0.25\) m. These conditions encode balance and contact constraints, while \(q_d\) captures bandwidth limits through velocity and acceleration tracking. On held-out data, the success head reaches AUROC \(0.979\) and AUPRC \(0.997\), with failure recall \(0.93\). Rank agreement with the simulator oracle is reported as within-prompt Kendall \(\tau=0.656\) overall and \(0.674\) on mixed success/failure prompts [2606.22998].

## 4. Semantic alignment verifier and motion generator

The semantic verifier is a learned text–motion co-embedding model trained directly on robot-skeleton motions. The motion encoder consists of a MovementConvEncoder with two temporal Conv1d downsampling layers followed by a BiGRU, yielding a 512-dimensional motion embedding \(\phi_{\text{motion}}(m)\). Root \(XY\) positions are replaced by frame-to-frame velocities to enforce translation invariance. The text encoder uses 300-dimensional word embeddings, 15-dimensional part-of-speech embeddings, and a BiGRU to produce a 512-dimensional text embedding \(\phi_{\text{text}}(t)\) [2606.22998].

Training uses an all-pairs margin loss
\[
\mathcal{L}_{\text{match}}=\frac{1}{B}\sum_i d_{ii}+\frac{1}{B(B-1)}\sum_{i\neq j}[\delta-d_{ij}]_+^2,
\quad
d_{ij}=\|\phi_{\text{text}}(t_i)-\phi_{\text{motion}}(m_j)\|_2,
\]
with margin \(\delta=2.0\). At test time,
\[
R_{\text{text}}(t,m)=\exp\big(-\|\phi_{\text{text}}(t)-\phi_{\text{motion}}(m)\|_2\big)\in(0,1].
\]
A motion-to-text retrieval sanity check on 9,116 held-out pairs with 32 distractors yields \(R@1=0.747\) and \(R@3=0.935\), compared with a random baseline of approximately \(1/32\) [2606.22998].

The default generator in the reported system is FSQ-GPT. Motions are represented at 50 Hz with 36-dimensional frames. An FSQ VAE with two stride-2 temporal downsamples compresses a sequence to \(L=T/4\) tokens. The FSQ code levels are \([3,3,3,3,3,2,2,2,2,2]\), giving \(K=7{,}776\) discrete codes. Reconstruction uses SmoothL1 losses on pose, velocity, and acceleration with weights \(1.0\) for root position, \(1.0\) for root quaternion, \(2.0\) for joint position, \(1.0\) for root velocity, \(1.0\) for joint velocity, \(0.5\) for root acceleration, and \(0.5\) for joint acceleration. The language model is Flan-T5-base, approximately \(220\)M parameters, with vocabulary extended by the motion codes and special tokens, trained with teacher-forced cross-entropy [2606.22998].

Training data combine AMASS with HumanML3D captions, retargeted to Unitree G1, together with the CLAW corpus, using an 8:1:1 train/validation/test split and 9,116 held-out prompts. Sampling uses ancestral decoding with temperature \(1.0\), top-\(k=50\), and top-\(p=0.95\). Because the verifiers operate on decoded trajectories rather than generator internals, TEXEDO transfers in plug-and-play fashion to unseen generators such as Kimodo [2606.22998].

## 5. Evaluation, scaling behavior, and deployment results

Evaluation uses Unitree G1 with the frozen SONIC controller. Dynamics are measured with success rate, MPJPE in millimeters, \(E_{\text{acc}}\) in mm/frame\(^2\), \(E_{\text{vel}}\) in mm/frame, and \(Q^*\). Semantic quality is measured with a VLM-as-Judge ensemble using GPT-5.5 and Gemini-2.5, with two rubrics and two frame samplings. The main comparison is between Base (\(N=1\)), \(R_{\text{dyn}}\)-only selection, \(R_{\text{text}}\)-only selection, full TEXEDO, and an Oracle upper bound based on simulator rollouts [2606.22998].

At \(N=32\) with FSQ-GPT, the Base system achieves VLM \(5.722\), Succ \(0.873\), MPJPE \(44.34\), \(E_{\text{acc}}=6.09\), \(E_{\text{vel}}=11.82\), and \(Q^*=0.829\). \(R_{\text{dyn}}\)-only improves executability to Succ \(0.990\), MPJPE \(38.15\), \(E_{\text{acc}}=3.30\), \(E_{\text{vel}}=5.91\), and \(Q^*=0.945\), but semantics drop to VLM \(4.924\). \(R_{\text{text}}\)-only improves semantics to VLM \(6.110\), with more modest dynamics gains: Succ \(0.885\), MPJPE \(42.97\), \(E_{\text{acc}}=4.90\), \(E_{\text{vel}}=10.03\), \(Q^*=0.847\). Full TEXEDO balances the two objectives: VLM \(6.054\), Succ \(0.984\), MPJPE \(39.09\), \(E_{\text{acc}}=4.26\), \(E_{\text{vel}}=7.78\), and \(Q^*=0.926\) [2606.22998].

Relative to Base, TEXEDO yields \(+0.112\) Succ, \(-5.25\) mm MPJPE, \(-1.83\) mm/frame\(^2\) \(E_{\text{acc}}\), \(-4.04\) mm/frame \(E_{\text{vel}}\), \(+0.097\) in \(Q^*\), and \(+0.332\) VLM. The reported interpretation is that TEXEDO preserves most of the semantic benefit of \(R_{\text{text}}\)-only while retaining most of the executability benefit of \(R_{\text{dyn}}\)-only.

The same pattern persists under transfer and distribution shift. For unseen Kimodo, Base at \(N=1\) obtains VLM \(4.823\), Succ \(0.937\), MPJPE \(38.62\), \(E_{\text{acc}}=1.84\), \(E_{\text{vel}}=5.33\), \(Q^*=0.918\); with full TEXEDO at \(N=32\), these become VLM \(5.381\), Succ \(0.954\), MPJPE \(37.49\), \(E_{\text{acc}}=1.70\), \(E_{\text{vel}}=4.97\), and \(Q^*=0.935\). On BONES-SEED out-of-distribution prompts, Base reports VLM \(4.116\), Succ \(0.860\), MPJPE \(36.22\), \(E_{\text{acc}}=8.36\), \(E_{\text{vel}}=15.71\), \(Q^*=0.804\), while TEXEDO reaches VLM \(4.401\), Succ \(0.980\), MPJPE \(26.64\), \(E_{\text{acc}}=3.65\), \(E_{\text{vel}}=3.88\), and \(Q^*=0.942\) [2606.22998].

Real-world deployment on a Unitree G1 reports 30/30 successful executions for 30 prompts at \(N=32\), with representative tracking averages MPJPE \(37.41\) mm, \(E_{\text{vel}}=4.35\) mm/frame, and \(E_{\text{acc}}=2.67\) mm/frame\(^2\). The reported qualitative coverage includes locomotion, gestures, and compositional prompt combinations [2606.22998].

## 6. Limitations, scope, and research significance

TEXEDO introduces a specific form of controller awareness, but its guarantees remain conditional on the learned verifier and the downstream tracker. The paper identifies a latency–quality trade-off controlled by \(N\), verifier dependence on the controller used to label rollouts, the possibility of rare but consequential false positives, weaker semantic generalization on out-of-distribution prompts, and the open-loop character of the current selection procedure. The dynamic verifier is controller-specific; swapping controllers requires relabeling via new rollouts, although the underlying motion corpora can be reused. Multi-robot portability is therefore not automatic [2606.22998].

The framework is also informative as a methodological statement. It does not align generator and tracker by joint retraining, nor does it fine-tune the controller to absorb generator errors. Instead, it keeps both components frozen and inserts a controller-grounded verifier at runtime. This places TEXEDO within a broader class of test-time scaling methods, but with an unusual reward structure in which safety-like feasibility is hard-gated and semantics are optimized only after feasibility has been established.

Within language-conditioned motion generation, this architecture addresses a common failure mode of kinematically plausible but dynamically untrackable outputs. The reported avoided errors include foot scuffing that leads to trips, excessive pelvis pitch or roll, overly rapid limb swings that saturate actuators, and contact timing mismatches. This suggests that TEXEDO’s main contribution is not a new generative prior, but an executable selection policy grounded in the dynamics of a particular robot–controller pair [2606.22998].

The resulting picture is that TEXEDO converts extra inference-time samples into a single reference trajectory that is both executable and semantically faithful. In the reported experiments, that conversion improves simulation metrics, transfers across generators, generalizes to out-of-distribution prompts, and carries through to real-world humanoid deployment without modifying either the generator or the controller.

Source: https://www.emergentmind.com/topics/texedo