- The paper presents a test-time scaling framework that integrates candidate sampling with dual-verifier checks for controller feasibility and semantic alignment.
- It employs a modular pipeline using dynamic feasibility and semantic alignment verifiers, achieving a 0.984 tracking success rate and robust performance on real hardware.
- The framework generalizes to unseen generators and out-of-distribution prompts, offering a deployable solution for language-conditioned humanoid motion without retraining.
TEXEDO: Controller-aware Test-Time Scaling for Language-conditioned Humanoid Motion Generation
Motivation and Problem Statement
Language-conditioned motion generation interfaces have rapidly evolved for humanoid robots, relying on pre-trained generators mapping text prompts to motion sequences. However, such generators are almost universally trained on retargeted human motion datasets, which transfer only kinematic priors and lack explicit grounding in the physical constraints and operational envelope of actual whole-body humanoid controllers. This results in a deployment gap: generated motions are semantically plausible but frequently violate balance, contact, actuation, or other robot-specific constraints, leading to frequent execution failures or poorly specified behavior.
TEXEDO addresses this gap by introducing a controller-aware, test-time scaling framework that leverages inference-time diversity latent in a frozen generator. Instead of deploying a single sampled motion, TEXEDO samples a pool of candidate motions for each prompt and uses grounded verifiers—evaluating both dynamic feasibility and semantic alignment—for selection. This process transforms a generic, controller-agnostic motion generator into a deployable system tightly coupled to the robot controller, without retraining either component.
Methodology
Modular Sampling and Verification Pipeline
TEXEDO implements a three-stage pipeline:
- TEXT: For a given language prompt, sample N candidate motions using a black-box text-conditioned generator (instantiated as FSQ-GPT in main experiments).
- SEE: Score each candidate using two grounded verifiers:
- Dynamic Feasibility Verifier: A temporal Transformer distilled from offline controller rollouts, predicting whether a candidate can be executed by the controller (balance, contact, actuation limits).
- Semantic Alignment Verifier: Contrastive co-embedding model trained directly on robot-skeleton motions, measuring the alignment of the motion with the prompt in a learned metric space.
- DO: Filter candidates by dynamic feasibility (hard constraint) and rerank the survivors by semantic score. If no candidate is predicted feasible, select the least-bad motion based on partial progress and tracking quality.
This asymmetric selection strategy—feasibility as a constraint, semantics as an objective within feasible set—is designed to minimize deployment risk, prioritizing stability and safety.
Generator and Verifier Details
- FSQ-GPT Generator: Discretizes continuous motions via Finite Scalar Quantization (FSQ), then fine-tunes an instruction-tuned Flan-T5-base model to generate motion token sequences from text.
- Dynamic Feasibility Verifier: Trained on composite oracle quality from offline SONIC rollouts, incorporates binary success, tracking quality, and progress ratio predictions. Achieves strong AUROC (0.979), high recall, and robust rank correlation (Kendall Ï„=0.656) with oracle metrics.
- Semantic Alignment Verifier: BiGRU encoders for text and motion, trained on all-pairs margin contrastive loss, yielding high retrieval recall in motion-to-text matching (R@1=0.747 with 32 distractors).
Experimental Results
- Intra-candidate selection: Dynamic feasibility and semantic alignment verifiers independently convert candidate diversity into improvements on execution quality and alignment, respectively. The full TEXEDO selector balances these competing axes, outperforming pure feasibility or semantics selectors when N=32.
- Numerical results: TEXEDO achieves $0.984$ tracking success rate (Succ), composite oracle quality (Q∗=0.926), and a VLM-Judge semantic score of $6.054$, validating superior joint performance over baselines. Extreme optimization of either axis degrades the other; TEXEDO's composition resolves this tension.
Generalization
- Zero-shot transfer: Verifiers trained on FSQ-GPT rollouts generalize to unseen generators such as Kimodo, improving both execution and alignment metrics without retraining.
- Out-of-distribution prompts: TEXEDO improves success rate and reduces tracking errors on BONES-SEED prompts that are outside the training distribution, with only minor semantic performance decay, demonstrating strong feasibility verifier transfer.
Real-world Deployment
- TEXEDO is deployed on a Unitree G1 humanoid. With N=32 candidates sampled per prompt, the framework successfully executes all 30 diverse prompts on physical hardware, spanning locomotion, upper-body gestures, and compound actions. Mean per-joint position error is $37.41$ mm, and all trajectories complete without falls or early termination.
Implications and Future Directions
The TEXEDO framework directly addresses the critical decoupling between motion generation and controller compatibility at deployment time. By treating controller feasibility as a non-negotiable constraint and leveraging sampling diversity for semantic fidelity, the system minimizes risk and maximizes practical utility. This paradigm is generator-agnostic and tracker-specific, with the verifiers operating on decoded motion trajectories rather than generator internals. This enables plug-and-play transfer across generators and robust adaptation to changed tracking controllers with minimal retraining, provided a reference motion corpus is available for relabeling feasibility targets.
Practically, TEXEDO unlocks reliable, language-guided robotic programming interfaces for complex humanoid platforms, where control risks are nontrivial and failure modes are safety-critical. The theoretical contribution is the compositionality between test-time diversity and post hoc verification, which may generalize to other multimodal generative models and complex downstream constraints in robotics.
Future developments could include adaptive sampling to further optimize latency–quality trade-offs, scalable verifier adaptation across tracker variants, and integration with even richer motion priors (including diffusion-based generators or explicit task constraints). Generating larger and more diverse candidate pools, potentially augmented by learned prioritization or preference models, could further enhance both semantic and physical fidelity at deployment.
Conclusion
TEXEDO presents a robust, controller-aware test-time scaling framework for language-conditioned humanoid motion generation, achieving strong motion executability and semantic fidelity without retraining the underlying generator or tracker. Its plug-in verifiers and compositional selection strategy enable generalization to unseen generators and out-of-distribution prompts, and its practical effectiveness is validated in real-world hardware deployment. TEXEDO advances the state-of-the-art in deployable, language-driven humanoid control by bridging the critical gap between generation and execution.