---
title: 'TLPO: Adaptive Multi-Expert Preference Optimization'
url: https://www.emergentmind.com/topics/timestep-layer-adaptive-multi-expert-preference-optimization-tlpo
type: topic
---

# TLPO: Adaptive Multi-Expert Preference Optimization

Timestep-Layer Adaptive Multi-Expert Preference Optimization (TLPO) is a two-stage training framework designed for aligning diffusion-based audio-driven portrait animation models with fine-grained, multidimensional human preferences. TLPO enables diffusion models to simultaneously optimize for motion naturalness, lip-sync accuracy, and visual quality by decoupling these potentially conflicting objectives into specialized expert modules that are adaptively fused across both denoising timesteps and transformer layers. This mechanism avoids the overfitting and mutual interference often observed when applying scalarized or undifferentiated reward optimization, leveraging the intrinsic stage-wise functional decomposition present in diffusion transformer (DiT) architectures [2508.11255].

## 1. Multidimensional Preference Alignment in Diffusion Models

Audio-driven portrait animation must meet several human-perceived criteria, primarily motion naturalness (MN), lip-sync accuracy (LS), and visual quality (VQ). These objectives are inherently competing; for example, maximizing lip-sync alignment may degrade motion fluidity or detail. Prior methods that collapse objectives into a single scalar reward have been prone to over-optimization of one dimension at the expense of others. Furthermore, diffusion transformers exhibit stepwise specialization—the initial denoising stages govern global structure (e.g., motion), while later steps refine details (e.g., facial texture), and different transformer layers attend to varying spatial-frequency content. Uniform preference injection fails to leverage this internal organization, motivating a more granular, stage-and-layer-aware preference modulation [2508.11255].

## 2. TLPO Framework and Architecture

TLPO is built upon a frozen DiT-based latent diffusion transformer backbone, pre-trained as Wan2.1, and augmented with a 3D variational autoencoder. Input audio features are extracted using Wav2Vec2 and injected via cross-attention mechanisms at every DiT block.

### Multi-Expert Structure

Three lightweight LoRA (Low-Rank Adaptation) expert modules are implemented in every linear sub-layer of each DiT block, each aligning with a single objective:
- $E_{m}$: Motion Naturalness expert
- $E_{l}$: Lip-Sync expert
- $E_{v}$: Visual Quality expert

Each expert LoRA injects a low-rank delta $\Delta h_{l}^{\,e} \in \mathbb{R}^{\text{hidden} \times r}$ to its respective frozen layer output $h_{l}$. For every inference pass, the final activation for layer $l$ at timestep $t$ is:
$$
h'_{l} = h_{l} + \sum_{e \in \{m, l, v\}} w_{e}^{\,l}(t)\, \Delta h_{l}^{\,e}
$$

### Timestep-Layer Adaptive Fusion

A gating network, parameterized by $\{W_{\rm gate}^l, \mathbf{b}^l\}$ for each layer $l$, takes as input the timestep embedding $t_{\rm emb} \in \mathbb{R}^d$ and outputs expert weights:
$$
\mathbf{w}^{\,l}(t) = \mathrm{softmax}(W_{\rm gate}^l\, t_{\rm emb}) + \mathbf{b}^l,
$$
where $W_{\rm gate}^l \in \mathbb{R}^{3 \times d}$, $\mathbf{b}^l \in \mathbb{R}^3$. This mechanism allows the model to dynamically select the degree of each expert's influence at each diffusion stage and transformer layer, with minimal computational overhead (less than 1% additional parameters) [2508.11255].

## 3. Training Strategy and Loss Functions

TLPO training proceeds in two distinct stages.

### Stage 1: Single-Expert DPO

Each expert module is independently optimized via Direct Preference Optimization (DPO), using pairs of samples $(x^+, x^-)$ curated for the target preference dimension and their respective reward function $R_e$. The loss is:
$$
\mathcal{L}_e = -\mathbb{E}_{(x^+,x^-)}\;\log \sigma\!\left( -\frac{\beta}{2}\bigl(L(x^+,t)-L(x^-,t)\bigr) \right)
$$
where $L(\cdot, t)$ is the denoising loss and $\beta$ is a temperature hyperparameter. For the lip-sync expert $E_{l}$, the loss is further reweighted via a lip-region mask $M$:
$$
\mathcal{L}_{l} = \sum_{p \in \text{pixels}} M(p)\,\mathcal{L}_\text{DPO}(x^+,x^-;p)
$$

### Stage 2: Fusion-Gate Optimization

All expert modules are frozen, and only the fusion gate parameters are updated using "full-dimension" pairs (real vs degraded synthetic samples) and a combined DPO loss:
$$
\mathcal{L}_\text{gate} = \sum_{e\in\{m,l,v\}} \sum_{(x^+,x^-)} -\log\sigma\!\left(-\frac{\beta'}{2}(L_{e}(x^+,t)-L_{e}(x^-,t))\right)
$$

This two-stage design ensures that gradients never push experts into direct competition and enables the fusion gate to modulate contributions from each expert without disrupting their optimized directions.

## 4. Data, Hyperparameters, and Implementation

Training leverages the Talking-NSQ dataset (410,000 auto-scored preference pairs: 180,000 for MN, 100,000 for LS, 130,000 for VQ) for expert adaptation, and 18,000 full-dimension pairs (real vs degraded) for gate fusion. LoRA modules use a rank $r=128$. Expert training employs an AdamW optimizer with LR = $1 \times 10^{-5}$, $\beta = 5000$, running MN/VQ for 10 epochs and LS for 20 epochs on 16$\times$A100 GPUs, while gate fusion uses LR = $1 \times 10^{-6}$, $\beta' = 1000$ for 5 epochs, updating only gating weights. The model runs for 50 diffusion timesteps, each with approximately 24 DiT blocks of 8 linear sub-layers each [2508.11255].

## 5. Performance and Empirical Analysis

Empirical evaluation demonstrates that TLPO surpasses four state-of-the-art baselines (FantasyTalking, HunyuanAvatar, OmniAvatar, MultiTalk) across all core metrics. The following table summarizes key results:

| Metric        | Baseline  | TLPO      |
|---------------|-----------|-----------|
| HKC (↑)       | 0.838     | 0.895     |
| Sync-C (↑)    | 3.154     | 5.704     |
| FID (↓)       | 43.137    | 35.438    |
| FVD (↓)       | 483.108   | 341.181   |

Ablation studies indicate that removing timestep gating, fusing at expert- or module-level, or using scalarized preference optimization significantly degrades performance, especially on motion and lip-sync dimensions. User ratings (on a 0–10 scale, 24 raters) show gains of 1.3 (MN), 0.8 (LS), and 1.0 (VQ) over the strongest baseline. Qualitative inspection reveals that TLPO yields more natural head/hand motion, accurate mouth shapes over long sequences, and sharper facial details compared to prior methods.

## 6. Limitations, Generalizations, and Future Directions

TLPO's two-stage training procedure introduces procedural complexity and necessitates careful curation of full-dimension preference pairs. While the gating mechanism adds minimal parameters, it can marginally increase inference latency. The framework, however, is broadly generalizable. The principal recipe—decoupling conflicting objectives into specialized expert adapters and dynamically reweighting them along network axes—may be extended to other multi-objective generative tasks, including:
- Text-to-image diffusion models balancing style and content
- Video style transfer, mediating temporal coherence and per-frame fidelity
- Any generative process where objectives are spatially, temporally, or functionally separable

Potential future directions noted include automatic discovery of new expert axes (e.g., emotion), meta-learning for adaptable gating policies, and unified, closed-loop optimization jointly training both reward model and generative process [2508.11255].

## 7. Summary

Timestep-Layer Adaptive Multi-Expert Preference Optimization (TLPO) enables diffusion-based generative models to resolve conflicts among multidimensional, possibly antagonistic, human preferences by isolating optimization processes and subsequently adaptively combining their influence at the level of diffusion timestep and network layer. Empirical results support TLPO's efficacy in aligning portrait animation outputs with human judgments on motion, lip synchronization, and visual quality, indicating its broader applicability for multi-objective preference alignment scenarios [2508.11255].

Source: https://www.emergentmind.com/topics/timestep-layer-adaptive-multi-expert-preference-optimization-tlpo