---
title: 'SoftCFG: Uncertainty-Guided AR Generation'
url: https://www.emergentmind.com/topics/softcfg
type: topic
---

# SoftCFG: Uncertainty-Guided AR Generation

to=arxiv__search _人人碰းjson  content='{"query":"arXiv:2510.00996 SoftCFG uncertainty-guided stable guidance for visual autoregressive model", "max_results": 5, "sort_by": "relevance"}'
to=arxiv__search  大发棋牌json  content='{"query":"\"SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model\"", "max_results": 10, "sort_by": "relevance"}'
to=arxiv__search  大发快三大小单双json  content='{"query":"SoftCFG uncertainty-guided stable guidance visual autoregressive", "max_results": 10, "sort_by": "relevance"}'
to=arxiv__search ացնումjson  content='{"query":"AliTok visual autoregressive ImageNet 256 2025", "max_results": 5, "sort_by": "relevance"}'
to=arxiv__search  微信天天彩票രണjson  content='{"query":"\"AliTok\" Wu 2025 autoregressive image generation", "max_results": 10, "sort_by": "relevance"}'
to=arxiv__search  海南天天中彩票json  content='{"query":"\"LuminaGPT2\" Xin 2025 autoregressive image model", "max_results": 10, "sort_by": "relevance"}'
SoftCFG is an inference-time modification of classifier-free guidance (CFG) for visual autoregressive (AR) image generators. It retains the standard conditional–unconditional combination used in CFG, but alters how the guidance signal is constructed in AR decoding by allowing every generated token in the sequence to contribute a certainty-weighted guidance signal through the model’s internal memory, specifically the unconditional value cache. The method was introduced in "SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model" [2510.00996]. Its stated purpose is to address two AR-specific pathologies—guidance diminishing, in which the conditional–unconditional gap rapidly vanishes as decoding progresses, and over-guidance, in which strong conditions distort visual coherence—while remaining training-free, model-agnostic, and compatible with existing AR pipelines. The paper reports that SoftCFG improves image quality over standard CFG and achieves state-of-the-art FID on ImageNet \(256\times 256\) among autoregressive models [2510.00996].

## 1. Formal setting and relation to standard CFG

SoftCFG is defined in the context of visual AR models that generate an image as a sequence of discrete tokens \((x_1,\dots,x_T)\) using a decoder-only transformer. Conditions such as class labels or text are introduced as special tokens at the beginning of the sequence. At decoding step \(t\), standard AR CFG runs the model twice: once with conditioning \(c\), and once with the conditioning token replaced by an empty token \(\emptyset\). With key/value caches for prior tokens, the two branches are

$$
\mathbf{z}_{t}^{\text{cond}} = f_\theta(\mathbf{K}_{<t}^{\text{cond}}, \mathbf{V}_{<t}^{\text{cond}}, x_{t-1}, c),
$$

$$
\mathbf{z}_{t}^{\text{uncond}} = f_\theta(\mathbf{K}_{<t}^{\text{uncond}}, \mathbf{V}_{<t}^{\text{uncond}}, x_{t-1}, \emptyset).
$$

The usual AR CFG rule combines these logits as

$$
\mathbf{z}_{t}^{\text{CFG}} = \mathbf{z}_{t}^{\text{uncond}} + \gamma\big(\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\big),
$$

with guidance scale \(\gamma>0\), and the sampling distribution

$$
p_\theta^{\text{CFG}}(\cdot \mid x_{<t}, c) = \mathrm{softmax}(\mathbf{z}_{t}^{\text{CFG}}).
$$

The paper denotes the guidance offset by \(\Delta_t=\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\). In current AR implementations, CFG effectively perturbs only the first class token, because the conditional and unconditional branches differ at the conditioning token while sharing the same growing context of visual tokens thereafter. SoftCFG preserves the two-branch setup and the guidance-scale parameter \(\gamma\), but shifts the perturbation mechanism from a purely logit-level global offset toward a context-aware perturbation of the unconditional branch’s value cache [2510.00996].

## 2. AR-specific pathologies motivating SoftCFG

The central motivation for SoftCFG is that CFG behaves differently in AR generation than in diffusion. Diffusion models re-inject conditioning at every denoising step, whereas AR models mainly condition via early tokens. As the generated visual context grows, the class or text token becomes increasingly distant from later predictions, and the conditional and unconditional branches become more similar. The paper terms this phenomenon guidance diminishing [2510.00996].

To quantify this effect, the paper measures entropy of the token distribution. For vocabulary size \(V\),

$$
p_t(i)=\frac{\exp(z_{t,i})}{\sum_{j=1}^{V}\exp(z_{t,j})}, \qquad i=1,\dots,V,
$$

$$
H(p_t)=-\sum_{i=1}^{V} p_t(i)\log p_t(i),
$$

and normalized entropy

$$
\hat H(p_t)=\frac{H(p_t)}{\log V}, \qquad \hat H(p_t)\in[0,1].
$$

On AliTok-XL, the paper plots normalized entropy over generation steps for baseline sampling, CFG, and their difference. The reported result is that the entropy difference quickly approaches \(1\), which is interpreted as the guidance effect becoming negligible even within short token grids such as \(16\times 16\). The consequences described are that later tokens are insufficiently influenced by the condition, text–image alignment deteriorates, and earlier errors are harder to correct.

The second pathology is over-guidance. When \(\gamma\) is too large, or when \(\|\Delta_t\|\) is large, the conditional branch dominates and can force semantics that conflict with the current visual context. The paper associates this with artifacts such as duplicated parts and unnatural shapes. Its illustrative example is a prompt emphasizing a “banana,” where strong guidance maps that attribute onto the elephant’s tusk, creating semantic misalignment. In LuminaGPT examples, standard CFG is described as producing tangled motorcycles, extra trunks, and redundant human hands, whereas SoftCFG is reported to reduce these artifacts. The paper explicitly likens the pair guidance diminishing versus over-guidance to gradient vanishing versus exploding in training, framing SoftCFG and Step Normalization as a regularization-and-normalization response to that duality [2510.00996].

## 3. Uncertainty-guided perturbation of the value cache

SoftCFG’s core idea is to let each generated token contribute certainty-weighted guidance through the model’s internal memory. For each previously generated token \(x_i\), the method uses the maximum probability from the conditional branch as a confidence proxy:

$$
p_{\max}(x_i)=\max_v\; p_\theta^{\text{cond}}(\cdot \mid x_{<i}, c),
$$

and defines

$$
w_i = 1-p_{\max}(x_i).
$$

Here \(w_i\) is described as an uncertainty weight. The paper states that tokens with higher confidence, corresponding to lower \(w_i\), are considered more reliable and thus receive stronger perturbations. The perturbation itself is applied to the unconditional value cache:

$$
\tilde{\mathbf{v}}_i^{\text{uncond, pertcontext}} = w_i\,\mathbf{v}_i^{\text{uncond}}, \qquad i<t.
$$

Using the perturbed unconditional cache \(\tilde{\mathbf{V}}_{<t}^{\text{uncond, pertcontext}}\), the model computes

$$
\tilde{\mathbf{z}}_{t}^{\text{uncond, pertcontext}} =
f_\theta(\mathbf{K}_{<t}^{\text{uncond}}, \tilde{\mathbf{V}}_{<t}^{\text{uncond, pertcontext}}, x_{t-1}, \emptyset).
$$

SoftCFG then forms guided logits as

$$
\mathbf{z}_{t}^{\text{SoftCFG}}=
\mathbf{z}_{t}^{\text{cond}}+
\gamma\big(\tilde{\mathbf{z}}_{t}^{\text{uncond, pertcontext}}-\mathbf{z}_{t}^{\text{cond}}\big),
$$

which the paper also rewrites as

$$
\mathbf{z}_{t}^{\text{SoftCFG}}=
\mathbf{z}_{t}^{\text{CFG}}+\gamma\,\Delta_t^{\text{context}},
\qquad
\Delta_t^{\text{context}}=
\tilde{\mathbf{z}}_{t}^{\text{uncond, pertcontext}}-\mathbf{z}_{t}^{\text{uncond}}.
$$

In Algorithm 1, the algebraically equivalent form is

$$
\mathbf{z}_{t}^{\text{SoftCFG}} \gets (1+\gamma)\mathbf{z}_{t}^{\text{cond}}-\gamma\,\tilde{\mathbf{z}}_{t}^{\text{uncond, pertcontext}}.
$$

The term \(\Delta_t^{\text{context}}\) is characterized as a context-aware regularizer. When it is aligned with the standard text-only guidance direction \(\Delta_t\), it amplifies effective guidance; when it is misaligned, it reduces the effective guidance magnitude and therefore counteracts over-guidance. Mechanistically, SoftCFG differs from standard CFG by touching every value-cache entry rather than only exploiting the conditional token at the beginning of the sequence. Because the value cache is reused at subsequent decoding steps through self-attention, the effect of confident tokens persists over time, embedding guidance into the AR model’s context memory rather than injecting it only as a per-step global logit offset [2510.00996].

The presentation includes a technical subtlety. Although \(w_i=1-p_{\max}(x_i)\) is termed an uncertainty weight and the cache scaling uses \(w_i\), the text also emphasizes that high-confidence tokens are the reliable guidance carriers. The paper’s algorithmic resolution is to use \(1-w_i\) as the quantity normalized by Step Normalization, so the normalized perturbation budget is distributed according to confidence rather than uncertainty.

## 4. Step Normalization and bounded long-sequence perturbation

SoftCFG alone introduces a cumulative-perturbation issue. Repeatedly scaling unconditional value vectors across a long sequence suppresses information in \(\mathbf{V}^{\text{uncond}}\); for high-confidence tokens, where \(w_i\to 0\), the suppression becomes strong. The paper formalizes the resulting deviation between SoftCFG and vanilla CFG as

$$
\Delta_t^{\text{context}}=
\tilde{\mathbf{z}}_{t}^{\text{uncond, pertcontext}}-\mathbf{z}_{t}^{\text{uncond}}.
$$

Assuming \(f_\theta\) is \(L_t\)-Lipschitz with respect to the value cache at step \(t\), the appendix gives the bound

$$
\|\Delta_t^{\text{context}}\|
\leq
L_t \cdot \sum_{i<t}(1-w_i)\|\mathbf{v}_i^{\text{uncond}}\|.
$$

Because \(\sum_{i<t}(1-w_i)\) can grow with sequence length, the deviation may explode as decoding proceeds. The paper associates this with instability, drift or collapse of unconditional logits, and poor sample diversity. In its entropy analysis, unnormalized SoftCFG shows exploding normalized entropy at later steps [2510.00996].

Step Normalization is introduced to remove the sequence-length dependence by normalizing the effective confidence contributions at each step:

$$
\hat w_i =
1-\frac{1-w_i}{\sum_{j=1}^{t-1}(1-w_j)},
\qquad
\sum_{i=1}^{t-1}(1-\hat w_i)=1.
$$

Algorithm 1 uses a small \(\varepsilon\) for numerical stability:

$$
\hat w_i \gets 1-\frac{1-w_i}{\sum_{j=1}^{t-1}(1-w_j)+\varepsilon}.
$$

Defining \(c_i=1-w_i=p_{\max}(x_i)\), Step Normalization yields

$$
c_i' = 1-\hat w_i = \frac{1-w_i}{\sum_j(1-w_j)},
\qquad
\sum_i c_i' = 1.
$$

The perturbed values then become

$$
\tilde{\mathbf{v}}_i^{\text{uncond}}=\hat w_i\,\mathbf{v}_i^{\text{uncond}}.
$$

The paper’s proposition states that under the same \(L_t\)-Lipschitz assumption,

$$
\|\Delta_t^{\text{context}}\|
=
\|\tilde{\mathbf{z}}_{t}^{\text{uncond, pert}}-\mathbf{z}_{t}^{\text{uncond}}\|
\leq
L_t \cdot \max_{i<t}\|\mathbf{v}_i^{\text{uncond}}\|.
$$

This replaces a bound that grows with \(\sum_i(1-w_i)\) by one that depends only on the maximum value-vector norm. In the paper’s interpretation, the total perturbation budget is fixed to one unit at every step, so adding more tokens redistributes the budget rather than increasing it. This is the principal stabilization mechanism for long sequences and is reported to keep normalized entropy stable over generation steps [2510.00996].

## 5. Sampling procedure and implementation characteristics

SoftCFG assumes a trained AR image model with a VQ tokenizer, a decoder-only transformer with key/value caches, and standard CFG support. It operates with conditional input \(c\), standard sampling hyperparameters such as temperature and top-\(k\)/top-\(p\), guidance scale \(\gamma\), and sequence length \(T\). The sampling loop described in Algorithm 1 is:

1. Initialize \(x_0\gets \langle bos\rangle\), empty conditional and unconditional caches, and a buffer for per-token confidence.
2. At each step \(t\), compute conditional logits \(\mathbf{z}_t^{\text{cond}}\).
3. For each previous token \(i<t\), retrieve \(p_{\max}(x_i)\) and form \(w_i=1-p_{\max}(x_i)\).
4. Apply Step Normalization to obtain \(\hat w_i\).
5. Rescale the unconditional value cache entrywise by \(\hat w_i\).
6. Compute perturbed unconditional logits \(\tilde{\mathbf{z}}_{t}^{\text{uncond, pertcontext}}\).
7. Form \(\mathbf{z}_{t}^{\text{SoftCFG}}\), sample \(x_t\sim \mathrm{Sample}(\mathrm{softmax}(\mathbf{z}_{t}^{\text{SoftCFG}}))\), and update both original caches and the stored max probability for the new token.
8. After \(T\) steps, decode the token sequence to pixels with the VQ decoder.

The method is characterized as training-free and model-agnostic because it introduces no architectural changes, no new layers, and no tokenizer or decoder modifications. It requires access only to logits, the value cache, and the conditioning mechanism already used for CFG. Its added cost is described as per-step \(\mathcal{O}(t)\) scalar operations on the value cache, which are negligible relative to attention and MLP computation. The paper further reports that inference speed is almost identical to standard CFG on AliTok-0.6B and LuminaGPT2-7B [2510.00996].

A practical detail emphasized in the hyperparameter discussion is that SoftCFG can be used with scheduled guidance. The paper gives

$$
\gamma_t=(\gamma-1)\cdot \tfrac{1}{2}\big(1-\cos((t/T)^k\pi)\big),
$$

where \(k\) controls how guidance is distributed over the sequence. On AliTok-XL, the reported optimal \(k\) is around \(1.5\)–\(1.6\). The paper also notes that because SoftCFG “naturally amplifies” guidance via context perturbation, effective \(\gamma\) values can be smaller than for vanilla CFG.

## 6. Empirical evaluation and qualitative behavior

The paper evaluates SoftCFG on class-conditional ImageNet-1K at \(256\times 256\), where 50k images are sampled on the validation label set, and on text-to-image benchmarks GenEval and DPG-Bench. The class-conditional experiments use AliTok (B, L, XL), described in the paper as a state-of-the-art AR model on ImageNet \(256\times 256\), and the text-to-image experiments use LuminaGPT2, described as a large AR image model with top performance on GenEval and DPG-Bench. Reported evaluation metrics for ImageNet include FID, Inception Score (IS), precision/recall, and sFID [2510.00996].

On the main ImageNet \(256\times 256\) comparison, the paper reports that AliTok-XL baseline achieves FID \(1.37\), IS \(321.4\), sFID \(7.29\), and recall \(0.64\), whereas AliTok-XL + SoftCFG achieves FID **1.27**, IS \(302.4\), sFID \(6.76\), and recall \(0.65\). Among AR models, this is presented as the best FID, surpassing RAR-XL at \(1.50\), MAR-H at \(1.55\), VAR-d30-re at \(1.73\), RandAR-XXL at \(2.15\), and LlamaGen-3B at \(2.18\). Relative to diffusion transformers, the paper lists DiT-XL/2 at \(2.27\), MDTV2-XL/2 at \(1.58\), and REPA-E at \(1.26\), and describes AliTok-XL + SoftCFG’s FID of \(1.27\) as statistically on par with state-of-the-art diffusion models.

The ablation study on AliTok-XL separates the contribution of SoftCFG and Step Normalization:

| Setting | FID | Other reported metrics |
|---|---:|---|
| Baseline (no CFG) | 1.76 | IS 221.2, sFID 5.55 |
| + CFG | 1.37 | IS 321.4, sFID 7.29 |
| + SoftCFG (no StepNorm) | 1.32 | IS 288.1, sFID 7.62 |
| + SoftCFG + StepNorm | 1.32 | IS 302.0, sFID 7.16 |
| + SoftCFG + StepNorm + Opt. | 1.27 | IS 302.4, sFID 6.70 |

The stated interpretation is that SoftCFG improves FID over vanilla CFG even without Step Normalization, while Step Normalization stabilizes IS and improves sFID. Joint tuning of \(\gamma\) and the cosine exponent \(k\) yields the best results.

Qualitatively, the paper attributes to SoftCFG better structural consistency and fewer artifacts under strong guidance. In LuminaGPT examples, SoftCFG is described as reducing tangled motorcycles, extra trunks, and extra hands. Confidence heatmaps align high-confidence regions with salient semantic parts such as faces and object cores, while ambiguous backgrounds receive lower confidence. The paper treats this as evidence that high-confidence tokens are effective guidance carriers. At the same time, the method exhibits a documented failure mode: for the prompt describing “A calico cat… patchwork of orange, black, and white fur… on top of a sleek Mercedes Benz…”, the patchwork pattern is applied to the car rather than the cat. The accompanying confidence map reportedly places high confidence on the car and low confidence on the cat, illustrating that SoftCFG depends on the quality and localization of the model’s confidence estimates [2510.00996].

## 7. Relation to adjacent guidance methods, limitations, and prospective extensions

Within the paper’s taxonomy, SoftCFG is positioned as an AR analogue of uncertainty- or semantics-aware guidance methods that modulate guidance internally rather than relying solely on a fixed global logit offset. The paper links it conceptually to semantic-aware CFG, token-level guidance for diffusion, self-attention guidance, perturbed-attention guidance, and smoothed energy guidance. Its stated novelty is that this philosophy is adapted to AR generation by operating on the value cache, which serves as the primary context memory in decoder-only AR models. The paper also contrasts SoftCFG with CCA, which seeks guidance-free AR sampling through retraining, and with architectural families such as RAR, VAR, and MAR, which explore AR conditioning strategies but do not focus on uncertainty-aware inference [2510.00996].

The limitations identified in the paper are specific. First, the method relies on confidence quality: it assumes that larger \(p_{\max}\) corresponds to better semantic consistency, yet AR confidence estimates are not always calibrated or well localized. The calico-cat failure case is presented as direct evidence. Second, Step Normalization may be rigid: because it enforces \(\sum_{i=1}^{t-1}(1-\hat w_i)=1\), the perturbation is capped to the equivalent of at most one token’s worth per step, which may under-utilize guidance in very long sequences. Third, although Step Normalization removes linear growth in the deviation bound, extremely long horizons such as ultra-high resolution generation or video may require more adaptive normalization. Fourth, the experimental scope is limited to class-conditional and text-to-image image generation; multimodal AR tasks, video synthesis, and RLHF-trained generators remain outside the reported evaluation.

The paper’s suggested future directions follow from those limitations. It proposes using external perceptual models such as DINOv3 for better semantic confidence estimation, exploring adaptive or hierarchical normalization that scales the perturbation budget with context complexity or length, and extending the underlying principle—that previously generated visual content can guide subsequent tokens—to other generative systems, including LLMs for text and AR video models. This suggests that SoftCFG is best understood not only as a particular AR image-sampling heuristic, but as a broader proposal for integrating internally derived confidence signals into guidance mechanisms while preserving the training-free structure of CFG.

Source: https://www.emergentmind.com/topics/softcfg