---
title: Ratio-Aware Adaptive Guidance Schedule
url: https://www.emergentmind.com/topics/ratio-aware-adaptive-guidance-schedule
type: topic
---

# Ratio-Aware Adaptive Guidance Schedule

Searching arXiv for recent papers on ratio-aware/adaptive guidance schedules and related CFG scheduling methods.
RATIO-aware adaptive guidance schedule denotes a family of step-dependent control rules that replace a fixed guidance coefficient with a quantity recomputed from the current relation between conditioning and prior dynamics. In the narrow sense, the term refers to RAAG, “Ratio Aware Adaptive Guidance,” a training-free schedule for flow-based generative models that dampens classifier-free guidance at early reverse steps when the relative strength of conditional to unconditional predictions spikes [2508.03442]. In a broader recent literature, closely related mechanisms adapt guidance from timestep-conditioned conditional–unconditional ratios in text-to-motion diffusion [2506.02452], prior-to-guidance energy ratios in geometry-aware diffusion guidance [2603.11509], signal-to-noise ratios in retrieval-augmented masked diffusion [2603.17677], and learned or uncertainty-conditioned schedules in diffusion and teacher–student control settings [2606.24025; 2605.26155].

## 1. Conceptual basis and departure from static guidance

The common point of departure is the standard fixed-scale guidance rule. In flow-based sampling, classifier-free guidance is written as
$$
v_{\text{cfg}}(x_t,c)=v_u(x_t)+w\big(v_c(x_t,c)-v_u(x_t)\big),
$$
with a constant guidance scale \(w\) [2508.03442]. In diffusion, the corresponding static formulation uses a fixed \(\omega\),
$$
G_s(\mathbf{x}_t, t, \mathbf{c}) = G_\theta(\mathbf{x}_t, t, \emptyset) + \omega \cdot \bigl(G_\theta(\mathbf{x}_t, t, \mathbf{c}) - G_\theta(\mathbf{x}_t, t, \emptyset)\bigr),
$$
or, in score notation, \(s_{\mathrm{CFG}} = s_0 + w \cdot \Delta s\) [2506.02452; 2603.11509]. These forms assume that the balance between semantic fidelity, generative freedom, or teacher influence is constant across the entire trajectory.

Recent work rejects that assumption for several distinct reasons. RAAG argues that fast low-step flow sampling is dominated by an early “RATIO spike,” making the earliest reverse steps acutely sensitive to guidance scale [2508.03442]. ANT argues that diffusion models recover low-frequency structure first and high-frequency details later, so textual semantics are most useful when coarse structure is still being formed [2506.02452]. MOG attributes high-scale CFG failures to a geometric mismatch, namely Euclidean extrapolation in ambient space that drives trajectories off the high-density data manifold [2603.11509]. ARAM argues that fixed guidance is brittle in retrieval-augmented masked diffusion because retrieved context can be reliable, irrelevant, or conflicting/noisy, and therefore should not be weighted uniformly across steps or tokens [2603.17677].

The resulting shift is from a single global hyperparameter to a trajectory-sensitive control law. Depending on the framework, the control signal is the relative norm of conditional and unconditional predictions, the ratio of prior energy to guidance energy, an SNR-like quantity, or an uncertainty proxy that is mapped into a bounded coefficient.

## 2. RAAG and the early-step RATIO spike in flow-based generation

RAAG formalizes the conditional–unconditional imbalance in rectified flow and flow-matching models through the velocity gap
$$
\delta(x_t,c)=v_c(x_t,c)-v_u(x_t),
$$
and the key diagnostic
$$
\mathrm{RATIO}(x_t,c)=\frac{\|\delta(x_t,c)\|_2}{\|v_u(x_t)\|_2}.
$$
At the final-noise end \(t\to 1\), the paper shows
$$
v_u(x_1)=x_1-\mu_u,\qquad v_c(x_1,c)=x_1-\mu_c,
$$
with
$$
\mu_u=\mathbb{E}[x_0],\qquad \mu_c=\mathbb{E}[x_0\mid c],
$$
so that
$$
\delta(x_1,c)=\mu_u-\mu_c,
$$
and therefore
$$
\mathrm{RATIO}_{t=1}(c)=\frac{\|\mu_c-\mu_u\|_2}{\|x_1-\mu_u\|_2}.
$$
The paper’s theoretical claim is that this spike is intrinsic to the data distribution, independent of model architecture, because it is tied to the conditional mean shift \(\mu_c-\mu_u\) rather than a specific network design [2508.03442].

RAAG further studies the separation \(A(t)=\|x(t)-y(t)\|_2\) between two guided trajectories and derives a Grönwall-type bound with an exponential term involving \(w\,p_{\max}\), leading to the conclusion that trajectory differences can grow roughly like
$$
A(t)\propto e^{\lambda\, w\,p_{\max} t},
$$
up to additive terms. The practical interpretation is that strong fixed guidance at the earliest step amplifies floating-point error, seed variation, stochasticity, model mismatch, and small early-step perturbations [2508.03442].

The adaptive rule is a closed-form exponential decay driven by the current RATIO:
$$
w(p)=1+(w_{\max}-1)\exp(-a p), \qquad p=\mathrm{RATIO}(x_t,c).
$$
When \(p\approx 0\), \(w(p)\approx w_{\max}\); when \(p\) is large, \(w(p)\to 1\). Because the RATIO is highest at the beginning and then decays, the effective guidance is conservative early and less damped later. Integration into a standard flow sampler is explicit: compute \(v_u\) and \(v_c\), form \(\delta\), evaluate \(p=\|\delta\|_2/\|v_u\|_2\), set \(w_t\), and use \(v_{\text{cfg}}=v_u+w_t\delta\). The method requires no retraining, no architectural change, no extra model forward passes, and negligible runtime overhead, and it is compatible with standard flow sampling and schedulers such as UniPC [2508.03442].

Empirically, RAAG is reported to enable up to \(3\times\) speedup on SD3.5, up to \(4\times\) speedup on Lumina-Next, and \(2\times\) faster sampling on WAN2.1-14B while preserving or improving quality, robustness, and semantic alignment. The paper also reports prompt-adherence gains on GenEval, including SD3.5 Single Object from \(96.25\%\) to \(98.75\%\) and Overall Score from \(0.9063\) to \(0.9188\), and on Lumina-Next Single Object from \(71.25\%\) to \(92.50\%\) and Overall Score from \(0.5313\) to \(0.6875\) [2508.03442].

## 3. Temporal-semantic ratio scheduling in text-to-motion diffusion

In ANT, the adaptive schedule appears as Dynamic Classifier-Free Guidance scheduling (DCFG), which is explicitly coupled to the Semantic Temporally Adaptive (STA) module. The central observation is that denoising is not uniform across time: early reverse steps need stronger semantic conditioning to establish the global motion structure, while later steps benefit more from weaker conditioning and efficient refinement of details. ANT therefore describes DCFG as a ratio-aware schedule that gradually shifts the model from a condition-heavy regime to a more unconditional regime as diffusion progresses [2506.02452].

ANT replaces the static guidance scale with a timestep-dependent \(\omega_t\):
$$
\omega_t=\omega_{\min}+\phi(t)(\omega_{\max}-\omega_{\min}),
$$
where \(\phi(t)\) is monotonically decreasing over timesteps. The instantiated cosine schedule is
$$
\omega _t = \max\left\{\omega _{\min } + \frac {1}{2}\left ( 1 + \cos \left(\lambda\frac {T-t}{T}\pi \right)\right )(\omega _{\max } - \omega _{\min }) , 0\right\}.
$$
In the implementation, \(T=50\), \(\lambda=1.5\), \(\omega_{\max}=3.0\), and \(\omega_{\min}=1.5\), with \(\omega_{\max}\) and \(\omega_{\min}\) selected by grid search on the validation set [2506.02452].

Using this schedule, the guided noise prediction becomes a time-varying interpolation between conditional and unconditional predictions. ANT goes further by introducing a hard switch to unconditional generation in the later denoising phase: “we omit the conditional branch altogether when \(t\) exceeds a certain threshold (e.g., \(t > 0.5T\)).” In that case,
$$
\hat{\epsilon}_t = \hat{\epsilon}_{\text{uncond}}.
$$
The practical recipe is: compute the timestep-specific text feature \(c_t \gets \mathrm{STA}(c,t)\), obtain \(\epsilon_{\text{cond}}\) and \(\epsilon_{\text{uncond}}\), compute \(\omega_t\), and use the guided combination before the threshold and unconditional prediction only afterward. In the reported setup, inference uses DPM-Solver with 10 actual sampling steps, and DCFG is described as “plug-and-play” for existing diffusion-based text-to-motion systems [2506.02452].

STA provides the semantic mechanism underlying DCFG. It modulates text features with timestep information via
$$
L_t = L \oplus z_t,
$$
followed by Adaptive Layer Normalization,
$$
\hat{L}_t = \gamma \cdot \frac{L_t + \alpha(z_t)}{\sigma + \varepsilon} + \beta,
$$
and then
$$
\hat{c}_t = \hat{L}_t \oplus\text{CrossAttention}(c, \hat{L}_t).
$$
The paper states that DCFG is built on top of the behavior induced by STA: because STA makes text attention naturally weaken over time, the guidance schedule follows the same trend, reducing the CFG scale as semantic reliance fades [2506.02452].

The reported trade-off is explicitly not accuracy-only. On HumanML3D, ANT (StableMoFusion) without DCFG reports FID \(0.071^{\pm0.004}\), Top-1 R-Precision \(0.565^{\pm0.006}\), Top-2 \(0.756^{\pm0.003}\), and Top-3 \(0.843^{\pm0.002}\), while full ANT (StableMoFusion) with DCFG reports FID \(0.099^{\pm0.004}\), Top-1 \(0.560^{\pm0.002}\), Top-2 \(0.751^{\pm0.003}\), and Top-3 \(0.841^{\pm0.002}\). The “w/o DCFG” row is thus slightly stronger on the reported accuracy metrics, but sampling time per batch drops from \(0.949\) s to \(0.741\) s when DCFG is enabled, which the paper describes as a \(21.9\%\) efficiency improvement with “almost no loss in accuracy” [2506.02452].

## 4. Geometry-aware and objective-driven schedule optimization

A distinct line of work recasts guidance scheduling as either a geometry-aware control problem or an explicit objective-optimization problem over the reverse trajectory. In MOG, the unconditional and conditional scores are
$$
s_0(x_t,t)\triangleq s_\theta(x_t,t,\varnothing), \qquad s_c(x_t,t)\triangleq s_\theta(x_t,t,c),
$$
with
$$
\Delta s(x_t,t)\triangleq s_c(x_t,t)-s_0(x_t,t).
$$
The method defines the local objective
$$
\mathcal{J}_t(u_t)=\frac{1}{2}u_t^\top \mathbf{M}_t u_t + \beta(t)\langle \nabla_x \mathcal{E},u_t\rangle,
$$
where \(\mathbf{M}_t\succ0\) is a Riemannian metric and \(\mathcal{E}(x_t,c)\triangleq -\log p_t(c\mid x_t)\). Minimization gives
$$
u_t^\star=-\beta(t)\mathbf{M}_t^{-1}\nabla_x \mathcal{E}(x_t,c),
$$
and therefore the geometry-aware guided score
$$
s_{\mathrm{MOG}}=s_0+\beta(t)\mathbf{M}_t^{-1}\Delta s.
$$
Auto-MOG chooses \(\beta(t)\) by balancing the metric norm of the guidance update against the metric norm of the unconditional score:
$$
\big\|\beta(t)\,v_{\mathrm{nat}}\big\|_{\mathbf{M}_t}=\gamma\,\|s_0\|_{\mathbf{M}_t},
\qquad
v_{\mathrm{nat}}=\mathbf{M}_t^{-1}\Delta s,
$$
which yields
$$
\beta_{\mathrm{auto}}(t)=\gamma \sqrt{\frac{s_0^\top \mathbf{M}_t s_0}{\Delta s^\top \mathbf{M}_t^{-1}\Delta s}}
$$
and the practical form
$$
\beta_{\mathrm{auto}}(t)=\gamma\cdot \frac{E_{\mathrm{prior}}(t)}{E_{\mathrm{guid}}(t)+\varepsilon}.
$$
The appendix further gives
$$
\beta(t)=\mathrm{clamp}\!\left(\gamma \frac{\|s_0\|_{\mathbf{M}_t}}{\|\mathbf{M}_t^{-1}\Delta s\|_{\mathbf{M}_t}+10^{-6}}, 0, 50\right).
$$
This makes the “ratio-aware” designation literal: the scale is proportional to the ratio of prior energy to guidance energy [2603.11509].

Auto-MOG is presented as eliminating the need for manual hyperparameter tuning, with \(\gamma=1\) across main experiments. On SD-XL, the paper reports FID \(21.60\), HPSv2 \(29.00\), CLIP \(34.20\), Saturation \(0.17\), and Contrast \(0.16\), and on FLUX.1 it reports FID \(17.84\) together with the best alignment metrics reported. In an SD-XL ablation, Auto-MOG is compared with tuned fixed CFG: CFG \(w=7.5\) yields FID \(25.90\), HPSv2 \(28.70\), CLIP \(34.10\), whereas Auto-MOG yields FID \(25.10\), HPSv2 \(29.36\), CLIP \(34.67\) [2603.11509].

The information-theoretic schedule-optimization framework takes a different route. It defines a clean endpoint reference
$$
q_{0,\mathrm{clean}}^\lambda(x_0\mid y)\propto p_0(x_0)p_0(y\mid x_0)^\lambda,
$$
and optimizes the actual sampler-induced endpoint distribution toward this reference by minimizing
$$
\mathbb E_{p(y)} \mathrm{KL}\!\left( q_0^{\mathbf w}(x_0\mid y)\,\|\, q_{0,\mathrm{clean}}^\lambda(x_0\mid y) \right).
$$
The objective decomposes into a consistency term and a coverage term:
$$
\mathcal L_{0}({\mathbf w};\lambda) := -\lambda \mathbb E_{p(y)q_0^{\mathbf w}(x_0\mid y)} \left[ \log p_0(y\mid x_0) \right] + \mathbb E_{p(y)} \mathrm{KL}\!\left( q_0^{\mathbf w}(x_0\mid y)\,\|\, p_0(x_0) \right).
$$
The paper derives trajectory-level identities for both terms and optimizes one \(w_k\) per sampler step using a discretized loss, clipped updates, and resampling-based acceptance. Practical details include 128 generated samples per schedule optimization, 2 Hutchinson noise vectors for divergence estimation, 15 optimization iterations, and about 15 minutes on a single B200 GPU per schedule [2606.24025].

The learned schedules are reported to be non-uniform: weak guidance at high noise, stronger guidance in selected middle-noise intervals, and selectively active guidance again at low noise. On ImageNet-512 with EDM-XXL, the best adaptive schedule reaches FID \(1.45\) at \(\bar w=1.2\), competitive with interval guidance’s \(1.41\) and better than constant guidance’s \(1.83\); at \(\bar w=2\), adaptive guidance reaches FID \(2.38\), better than constant guidance’s \(5.54\) and interval guidance’s \(3.53\). On COCO with SD-XL, adaptive schedules achieve the highest CLIP score in all matched-\(\bar w\) groups and the best FID at \(\bar w=7\) and \(\bar w=9\) [2606.24025].

## 5. Retrieval-augmented and partially observable settings

In retrieval-augmented masked diffusion, ARAM defines ratio-aware adaptation at the level of individual masked tokens and denoising steps. For a masked token \(x\) at step \(t\), the model computes the retrieval-conditioned distribution
$$
p_{\text{cond}}(x)=p_\theta(x\mid x_t,q,\mathcal C)
$$
and the retrieval-free prior
$$
p_{\text{prior}}(x)=p_\theta(x\mid x_t,q).
$$
Standard discrete CFG uses
$$
\ell_{\lambda}=\ell_{\text{prior}}+\lambda(\ell_{\text{cond}}-\ell_{\text{prior}}),
$$
but ARAM replaces the fixed \(\lambda\) with an adaptive \(\lambda_t\) derived from an SNR-style ratio [2603.17677].

The paper defines a context score
$$
s(x;\mathcal C):=\log \frac{p_{\text{cond}}(x)}{p_{\text{prior}}(x)},
$$
and a retrieval information gain
$$
\mathrm{IG}_t:=D_{\mathrm{KL}}(p_{\text{cond}}\|p_{\text{prior}})=\mathbb E_{p_{\text{cond}}}[s(x;\mathcal C)].
$$
A local Taylor expansion around \(\lambda=0\) motivates
$$
\lambda^*=\frac{\mathbb{E}_{p_{\text{cond}}}[s]-\mathbb{E}_{p_{\text{prior}}}[s]}{\mathrm{Var}_{p_{\text{prior}}}(s)},
$$
which the paper interprets as
$$
\lambda^*=\frac{\text{Signal}}{\text{Noise}}.
$$
The practical algorithm replaces the variance term with entropy. The signal proxy is the symmetrized KL divergence
$$
D_{\mathrm{KL}}(p_{\text{cond}}\|p_{\text{prior}})+D_{\mathrm{KL}}(p_{\text{prior}}\|p_{\text{cond}}),
$$
the noise proxy is the conditional entropy
$$
\mathcal H(p_{\text{cond}})=-\sum_{x\in\mathcal V} p_{\text{cond}}(x)\log p_{\text{cond}}(x),
$$
and the adaptive weight is
$$
\lambda_t=\lambda_{\max}\cdot \tanh\left(\beta\cdot \frac{\text{Signal}}{\text{Noise}+\epsilon}\right).
$$
The method is explicitly training-free, step-wise and token-wise, and the paper reports that entropy works better than variance-based noise while \(\tanh\) improves stability and performance over raw SNR scaling [2603.17677].

In teacher–student reinforcement learning under partial observability, BA-GSAC studies adaptive guidance in a different formal setting. Guided SAC uses the control objective
$$
J_c=\mathbb{E}_{h_t,a \sim \pi_c}\left[Q(s,a)-\alpha \log \pi_c(a|h_t)\right]-\lambda \|\pi_c(h_t)-D(h_t)\|^2,
$$
where \(\lambda\) is the distillation coefficient from a privileged full-state teacher to a partial-observation student. BA-GSAC replaces fixed \(\lambda\) with a state-dependent \(\lambda_t\) derived from ensemble disagreement:
$$
u(h_t,a_t)=\frac{2}{N(N-1)}\sum_{i<j}\|\hat{o}_{t+1}^{(i)}-\hat{o}_{t+1}^{(j)}\|^2,
$$
$$
\lambda_t=\lambda_{\min}+(\lambda_{\max}-\lambda_{\min})\cdot \sigma\!\left(\frac{u_t-u_{\text{lo}}}{u_{\text{hi}}-u_{\text{lo}}}\right),
\qquad
\sigma(x)=\mathrm{clip}(x,0,1).
$$
Warmup calibration uses \(W=800\), \(\lambda_{\min}=0.01\), \(\lambda_{\max}=0.5\), and ensemble size \(N=5\). The paper also tests a deterministic linear decay baseline,
$$
\lambda_t=\lambda_{\max}(1-t/T)+\lambda_{\min}(t/T),
$$
with the same bounds \([0.01,0.5]\) [2605.26155].

The empirical findings are explicitly nuanced. Under mild and moderate partial observability, preliminary single-seed runs suggest benefits for adaptive guidance. Under severe occlusion, evaluated with 3 seeds for all methods, the adaptive coefficient collapses to \(\lambda_{\min}\) within about 3K steps. The severe-POMDP results are: Vanilla SAC mean \(75.2\), CV \(16.8\%\); fixed \(\lambda=0.01\) mean \(96.2\), CV \(29.8\%\); GSAC \((\lambda=0.1)\) mean \(105.6\), CV \(22.2\%\); BA-GSAC mean \(89.2\), CV \(13.3\%\); linear decay mean \(116.5\), CV \(8.9\%\). The paper attributes BA-GSAC’s failure to “observability blindness”: because the ensemble predicts partial observations, it achieves low disagreement even under heavy occlusion, modeling what is visible but unable to detect what is missing. A proposed architectural fix is to train the ensemble on full-state predictions using the guiding actor’s privileged access, but this fix is not validated there [2605.26155].

## 6. Comparative interpretation, limitations, and recurrent patterns

Across these methods, the diagnostic quantity being measured is not uniform, but it is always used to modulate guidance relative to an internal baseline rather than to hold guidance constant.

| Method | Measured quantity | Guidance rule |
|---|---|---|
| RAAG | \(\mathrm{RATIO}=\|\delta\|_2/\|v_u\|_2\) | \(w(p)=1+(w_{\max}-1)e^{-ap}\) |
| ANT / DCFG | timestep-dependent conditional–unconditional contribution | cosine-decayed \(\omega_t\), then unconditional-only after threshold |
| Auto-MOG | prior energy / guidance energy | \(\beta_t=\gamma \frac{\|s_0\|_{\mathbf{M}_t}}{\|\mathbf{M}_t^{-1}\Delta s\|_{\mathbf{M}_t}+\varepsilon}\) |
| ARAM | signal / noise | \(\lambda_t=\lambda_{\max}\tanh\!\left(\beta \frac{\text{Signal}}{\text{Noise}+\epsilon}\right)\) |
| BA-GSAC | disagreement-mapped uncertainty | clipped linear map into \([\lambda_{\min},\lambda_{\max}]\) |
| Information-theoretic CFG | trajectory objective for consistency–coverage | learned per-step \(w_k\) |

A recurring misconception is that adaptive guidance should have a universal schedule shape. The reported papers do not support that view. RAAG damps guidance at the earliest reverse steps because the first reverse step has the largest RATIO and strong guidance there causes exponential error amplification [2508.03442]. ANT instead argues for stronger guidance early and weaker guidance late because textual semantics are most useful while coarse motion structure is being established, after which conditional computation becomes increasingly wasteful [2506.02452]. The information-theoretic method learns weak high-noise guidance, stronger middle-noise guidance, and selective low-noise guidance rather than a monotone decay [2606.24025]. This suggests that the useful schedule shape depends on which failure mode is being controlled: early instability, temporal-semantic mismatch, manifold drift, retrieval conflict, or uncertainty miscalibration.

Another recurrent limitation is that adaptive guidance is not uniformly superior on every metric. ANT reports that the “w/o DCFG” row is slightly stronger on the cited HumanML3D accuracy metrics, even though DCFG improves efficiency by \(21.9\%\) with “almost no loss in accuracy” [2506.02452]. BA-GSAC shows a stronger counterexample: in severe POMDPs, a simple deterministic linear decay achieves the best performance across all metrics, while the uncertainty-reactive coefficient collapses early because the ensemble cannot detect missing information [2605.26155]. The information-theoretic optimizer likewise notes that exact gradient computation with respect to \(\mathbf w\) is hard because changing one weight changes future trajectory distributions, so the optimization uses a fixed-trajectory proposal direction plus resampling-based acceptance and remains an approximation to full trajectory optimization [2606.24025].

A broader implication is that ratio-aware adaptation is best understood as a control principle rather than a single algorithmic template. Some methods are explicitly training-free and plug-and-play, such as RAAG and ARAM [2508.03442; 2603.17677]. Some rely on architectural coupling, such as ANT’s dependence on STA-induced timestep-conditioned text features [2506.02452]. Some replace Euclidean extrapolation with metric-preconditioned guidance, as in MOG [2603.11509]. Others optimize a schedule against a clean endpoint reference rather than deriving it from a closed-form ratio at runtime [2606.24025]. A plausible implication is that future progress will depend less on finding one universally optimal schedule and more on matching the diagnostic ratio to the specific source of guidance failure in each model class and application regime.

Source: https://www.emergentmind.com/topics/ratio-aware-adaptive-guidance-schedule