TEXEDO: Controller-Aware Humanoid Motion
- TEXEDO is a controller-aware test-time scaling framework that transforms a frozen text-conditioned generator into a deployment-ready system via candidate sampling and dynamic feasibility verification.
- It employs dual verifiers—one enforcing a hard dynamic feasibility gate and another ensuring semantic alignment—to select motions that are both safe and faithful to the input text.
- The best-of-N selection process, validated on Unitree G1, significantly improves success rates and motion quality, making it a practical approach for language-conditioned humanoid robotics.
Searching arXiv for TEXEDO and closely related humanoid motion generation papers. Search query: TEXEDO controller-aware language-conditioned humanoid motion generation arXiv TEXEDO is a controller-aware test-time scaling framework for language-conditioned humanoid motion generation. It upgrades a frozen text-conditioned generator into a deployment-oriented system by sampling multiple candidate motions for a prompt and selecting the candidate that is both physically executable by a specific whole-body tracking controller and semantically aligned with the text. In the reported instantiation, TEXEDO operates on Unitree G1 whole-body trajectories, uses a dynamic feasibility verifier distilled from controller rollouts together with a semantic alignment verifier in a learned text–motion co-embedding space, and enforces a hard feasibility gate before semantic re-ranking (Cao et al., 22 Jun 2026).
1. Problem setting and design rationale
Language-conditioned whole-body motion generation is motivated by the prospect of using text as an interface for programming humanoid robots. The central difficulty addressed by TEXEDO is that most generators are trained on human motion corpora retargeted to robot morphologies, whereas deployment is governed by a particular whole-body tracking controller whose executable envelope depends on balance, contact timing, actuation limits, and controller-specific failure modes. A generated motion can therefore be semantically plausible while remaining difficult or impossible for the robot to execute (Cao et al., 22 Jun 2026).
The framework is organized as “Text-See-Do.” In the “TEXT” stage, a pretrained text-conditioned generator maps a prompt to a set of sampled candidate motions
where . For Unitree G1, the state dimension is , comprising root position, root quaternion, and 29 joints at $50$ Hz. In the “SEE” stage, each candidate is scored by a dynamic feasibility verifier and a semantic alignment verifier . In the “DO” stage, dynamic feasibility is treated as a hard constraint and semantic alignment as the objective: The deployed form is
The reported threshold is 0, chosen on a validation split to minimize false positives, and if no candidate passes the gate the method falls back to 1. This fallback is reported to occur for only 2 of prompts (Cao et al., 22 Jun 2026).
This design reframes motion generation as a selection problem rather than a retraining problem. A plausible implication is that TEXEDO is especially useful when the generator and the controller are already available but mismatched, because the method leaves both frozen and inserts grounded verification at inference time.
2. Controller-aware test-time scaling pipeline
TEXEDO’s defining operation is best-of-3 selection under grounded constraints. As 4 increases, the system exploits the diversity of the generator’s output distribution rather than attempting to alter that distribution. The paper reports a monotonic test-time scaling law: using the appropriate verifier, larger candidate pools improve executability and semantic quality without requiring a stronger generator (Cao et al., 22 Jun 2026).
The practical selection algorithm is straightforward. For a prompt 5, the system samples 6 independent motions, decodes them to continuous trajectories, computes 7 and 8, filters candidates with 9, and selects the feasible candidate with maximal 0. If the feasible set is empty, it returns the dynamically safest motion according to 1. This produces a conservative failure mode: semantics are sacrificed only when no candidate satisfies the feasibility gate.
The runtime cost is reported for an A100 40GB GPU with batched sampling and scoring. Sampling 2 candidates takes 3 s, scoring 4 candidates takes approximately 5 ms for 6 and 7 ms for 8, and the end-to-end latency for 9 is approximately 0 s per prompt. These numbers define the practical operating point used throughout the main experiments (Cao et al., 22 Jun 2026).
A central conceptual feature is the asymmetry between the two verifiers. Dynamic feasibility is not used as a soft preference; it is a safety gate. Semantic alignment is not used to rescue unsafe trajectories; it only ranks candidates inside the feasible set. This sharply separates “can the robot execute it?” from “does it match the prompt?”
3. Dynamic feasibility verifier
The dynamic feasibility verifier is distilled from whole-body tracking rollouts. Reference motions are executed by a frozen SONIC whole-body tracker in MuJoCo until clip completion or early termination. Each rollout yields three supervisory signals: success 1, progress ratio 2, and normalized tracking quality 3. The quality term is
4
where the velocity and acceleration errors are normalized by 95th-percentile constants (Cao et al., 22 Jun 2026).
These rollout signals are collapsed into an oracle quality
5
with 6 and 7. The constraint
8
guarantees success-first ordering: every successful rollout outranks every failed rollout, while failures remain graded by progress and quality. TEXEDO’s predicted dynamic score is
9
The verifier input at each frame is a 94-dimensional feature vector: 36-dimensional raw pose together with first and second finite differences, specifically root dynamics 7, joint position 29, joint velocity 29, and joint acceleration 29. The temporal encoder is a 4-layer causal Transformer with 0, 4 heads, pre-LayerNorm, and mean-attention pooling, followed by three heads predicting 1, 2, and 3. Training uses
4
with 5 and 6 (Cao et al., 22 Jun 2026).
Executability is grounded in controller-aware termination criteria: anchor height deviation 7 with 8 m, anchor tilt exceeding 9, and end-effector clearance exceeding $50$0 m. These conditions encode balance and contact constraints, while $50$1 captures bandwidth limits through velocity and acceleration tracking. On held-out data, the success head reaches AUROC $50$2 and AUPRC $50$3, with failure recall $50$4. Rank agreement with the simulator oracle is reported as within-prompt Kendall $50$5 overall and $50$6 on mixed success/failure prompts (Cao et al., 22 Jun 2026).
4. Semantic alignment verifier and motion generator
The semantic verifier is a learned text–motion co-embedding model trained directly on robot-skeleton motions. The motion encoder consists of a MovementConvEncoder with two temporal Conv1d downsampling layers followed by a BiGRU, yielding a 512-dimensional motion embedding $50$7. Root $50$8 positions are replaced by frame-to-frame velocities to enforce translation invariance. The text encoder uses 300-dimensional word embeddings, 15-dimensional part-of-speech embeddings, and a BiGRU to produce a 512-dimensional text embedding $50$9 (Cao et al., 22 Jun 2026).
Training uses an all-pairs margin loss
0
with margin 1. At test time,
2
A motion-to-text retrieval sanity check on 9,116 held-out pairs with 32 distractors yields 3 and 4, compared with a random baseline of approximately 5 (Cao et al., 22 Jun 2026).
The default generator in the reported system is FSQ-GPT. Motions are represented at 50 Hz with 36-dimensional frames. An FSQ VAE with two stride-2 temporal downsamples compresses a sequence to 6 tokens. The FSQ code levels are 7, giving 8 discrete codes. Reconstruction uses SmoothL1 losses on pose, velocity, and acceleration with weights 9 for root position, 0 for root quaternion, 1 for joint position, 2 for root velocity, 3 for joint velocity, 4 for root acceleration, and 5 for joint acceleration. The LLM is Flan-T5-base, approximately 6M parameters, with vocabulary extended by the motion codes and special tokens, trained with teacher-forced cross-entropy (Cao et al., 22 Jun 2026).
Training data combine AMASS with HumanML3D captions, retargeted to Unitree G1, together with the CLAW corpus, using an 8:1:1 train/validation/test split and 9,116 held-out prompts. Sampling uses ancestral decoding with temperature 7, top-8, and top-9. Because the verifiers operate on decoded trajectories rather than generator internals, TEXEDO transfers in plug-and-play fashion to unseen generators such as Kimodo (Cao et al., 22 Jun 2026).
5. Evaluation, scaling behavior, and deployment results
Evaluation uses Unitree G1 with the frozen SONIC controller. Dynamics are measured with success rate, MPJPE in millimeters, 0 in mm/frame1, 2 in mm/frame, and 3. Semantic quality is measured with a VLM-as-Judge ensemble using GPT-5.5 and Gemini-2.5, with two rubrics and two frame samplings. The main comparison is between Base (4), 5-only selection, 6-only selection, full TEXEDO, and an Oracle upper bound based on simulator rollouts (Cao et al., 22 Jun 2026).
At 7 with FSQ-GPT, the Base system achieves VLM 8, Succ 9, MPJPE 0, 1, 2, and 3. 4-only improves executability to Succ 5, MPJPE 6, 7, 8, and 9, but semantics drop to VLM 00. 01-only improves semantics to VLM 02, with more modest dynamics gains: Succ 03, MPJPE 04, 05, 06, 07. Full TEXEDO balances the two objectives: VLM 08, Succ 09, MPJPE 10, 11, 12, and 13 (Cao et al., 22 Jun 2026).
Relative to Base, TEXEDO yields 14 Succ, 15 mm MPJPE, 16 mm/frame17 18, 19 mm/frame 20, 21 in 22, and 23 VLM. The reported interpretation is that TEXEDO preserves most of the semantic benefit of 24-only while retaining most of the executability benefit of 25-only.
The same pattern persists under transfer and distribution shift. For unseen Kimodo, Base at 26 obtains VLM 27, Succ 28, MPJPE 29, 30, 31, 32; with full TEXEDO at 33, these become VLM 34, Succ 35, MPJPE 36, 37, 38, and 39. On BONES-SEED out-of-distribution prompts, Base reports VLM 40, Succ 41, MPJPE 42, 43, 44, 45, while TEXEDO reaches VLM 46, Succ 47, MPJPE 48, 49, 50, and 51 (Cao et al., 22 Jun 2026).
Real-world deployment on a Unitree G1 reports 30/30 successful executions for 30 prompts at 52, with representative tracking averages MPJPE 53 mm, 54 mm/frame, and 55 mm/frame56. The reported qualitative coverage includes locomotion, gestures, and compositional prompt combinations (Cao et al., 22 Jun 2026).
6. Limitations, scope, and research significance
TEXEDO introduces a specific form of controller awareness, but its guarantees remain conditional on the learned verifier and the downstream tracker. The paper identifies a latency–quality trade-off controlled by 57, verifier dependence on the controller used to label rollouts, the possibility of rare but consequential false positives, weaker semantic generalization on out-of-distribution prompts, and the open-loop character of the current selection procedure. The dynamic verifier is controller-specific; swapping controllers requires relabeling via new rollouts, although the underlying motion corpora can be reused. Multi-robot portability is therefore not automatic (Cao et al., 22 Jun 2026).
The framework is also informative as a methodological statement. It does not align generator and tracker by joint retraining, nor does it fine-tune the controller to absorb generator errors. Instead, it keeps both components frozen and inserts a controller-grounded verifier at runtime. This places TEXEDO within a broader class of test-time scaling methods, but with an unusual reward structure in which safety-like feasibility is hard-gated and semantics are optimized only after feasibility has been established.
Within language-conditioned motion generation, this architecture addresses a common failure mode of kinematically plausible but dynamically untrackable outputs. The reported avoided errors include foot scuffing that leads to trips, excessive pelvis pitch or roll, overly rapid limb swings that saturate actuators, and contact timing mismatches. This suggests that TEXEDO’s main contribution is not a new generative prior, but an executable selection policy grounded in the dynamics of a particular robot–controller pair (Cao et al., 22 Jun 2026).
The resulting picture is that TEXEDO converts extra inference-time samples into a single reference trajectory that is both executable and semantically faithful. In the reported experiments, that conversion improves simulation metrics, transfers across generators, generalizes to out-of-distribution prompts, and carries through to real-world humanoid deployment without modifying either the generator or the controller.