Papers
Topics
Authors
Recent
Search
2000 character limit reached

SoftCFG: Uncertainty-Guided AR Generation

Updated 14 July 2026
  • SoftCFG is an autoregressive guidance method that integrates token-level uncertainty into the value cache to enhance the conditional-unconditional signal.
  • It employs uncertainty weights and step normalization to address issues like guidance diminishing and over-guidance, ensuring improved visual coherence.
  • Empirical evaluations on ImageNet demonstrate that SoftCFG achieves state-of-the-art FID without requiring additional training or model modifications.

to=arxiv_search _人人碰းjson content='{"7query7 SoftCFG uncertainty-guided stable guidance for visual autoregressive model7", "7max_results7 7query7, "7sort_by7 "7relevance7 to=arxivsearch 大发棋牌json content='{"7query7 Uncertainty-guided Stable Guidance for Visual Autoregressive Model7\7 "7max_results7 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7query7, "7sort_by7 "7relevance7 to=arxivsearch 大发快三大小单双json content='{"7query7 uncertainty-guided stable guidance visual autoregressive", "7max_results7 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7query7, "7sort_by7 "7relevance7 to=arxivsearch ացնումjson content='{"7query7 visual autoregressive ImageNet 7max_results7query7\7^ 7max_results7query7max_results7query7", "7max_results7 7query7, "7sort_by7 "7relevance7 to=arxivsearch 微信天天彩票രണjson content='{"7query7 Wu 7max_results7query7max_results7query7^ autoregressive image generation", "7max_results7 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7query7, "7sort_by7 "7relevance7 to=arxivsearch 海南天天中彩票json content='{"7query7 Xin 7max_results7query7max_results7query7^ autoregressive image model", "7max_results7 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7query7, "7sort_by7 "7relevance7 SoftCFG is an inference-time modification of classifier-free guidance (CFG) for visual autoregressive (AR) image generators. It retains the standard conditional–unconditional combination used in CFG, but alters how the guidance signal is constructed in AR decoding by allowing every generated token in the sequence to contribute a certainty-weighted guidance signal through the model’s internal memory, specifically the unconditional value cache. The method was introduced in "SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model" (&&&7query7&&&). Its stated purpose is to address two AR-specific pathologies—guidance diminishing, in which the conditional–unconditional gap rapidly vanishes as decoding progresses, and over-guidance, in which strong conditions distort visual coherence—while remaining training-free, model-agnostic, and compatible with existing AR pipelines. The paper reports that SoftCFG improves image quality over standard CFG and achieves state-of-the-art FID on ImageNet PRESERVED_PLACEHOLDER7query7^ among autoregressive models (&&&7query7&&&).

SoftCFG is defined in the context of visual AR models that generate an image as a sequence of discrete tokens PRESERVED_PLACEHOLDER_7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7^ using a decoder-only transformer. Conditions such as class labels or text are introduced as special tokens at the beginning of the sequence. At decoding step PRESERVED_PLACEHOLDER_7max_results7, standard AR CFG runs the model twice: once with conditioning PRESERVED_PLACEHOLDER_7sort_by7, and once with the conditioning token replaced by an empty token PRESERVED_PLACEHOLDER_7relevance7. With key/value caches for prior tokens, the two branches are

PRESERVED_PLACEHOLDER_7query7^

PRESERVED_PLACEHOLDER_7\7^

The usual AR CFG rule combines these logits as

ztCFG=ztuncond+γ(ztcond−ztuncond),\mathbf{z}_{t}^{\text{CFG}} = \mathbf{z}_{t}^{\text{uncond}} + \gamma\big(\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\big),

with guidance scale γ>0\gamma>0, and the sampling distribution

pθCFG(⋅∣x<t,c)=softmax(ztCFG).p_\theta^{\text{CFG}}(\cdot \mid x_{<t}, c) = \mathrm{softmax}(\mathbf{z}_{t}^{\text{CFG}}).

The paper denotes the guidance offset by PRESERVED_PLACEHOLDER_7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7query7. In current AR implementations, CFG effectively perturbs only the first class token, because the conditional and unconditional branches differ at the conditioning token while sharing the same growing context of visual tokens thereafter. SoftCFG preserves the two-branch setup and the guidance-scale parameter PRESERVED_PLACEHOLDER_7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7, but shifts the perturbation mechanism from a purely logit-level global offset toward a context-aware perturbation of the unconditional branch’s value cache (&&&7query7&&&).

7max_results7. AR-specific pathologies motivating SoftCFG

The central motivation for SoftCFG is that CFG behaves differently in AR generation than in diffusion. Diffusion models re-inject conditioning at every denoising step, whereas AR models mainly condition via early tokens. As the generated visual context grows, the class or text token becomes increasingly distant from later predictions, and the conditional and unconditional branches become more similar. The paper terms this phenomenon guidance diminishing (&&&7query7&&&).

To quantify this effect, the paper measures entropy of the token distribution. For vocabulary size PRESERVED_PLACEHOLDER_7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7max_results7,

PRESERVED_PLACEHOLDER_7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7sort_by7^

PRESERVED_PLACEHOLDER_7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7relevance7^

and normalized entropy

PRESERVED_PLACEHOLDER_7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7query7^

On AliTok-XL, the paper plots normalized entropy over generation steps for baseline sampling, CFG, and their difference. The reported result is that the entropy difference quickly approaches PRESERVED_PLACEHOLDER_7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7\7, which is interpreted as the guidance effect becoming negligible even within short token grids such as PRESERVED_PLACEHOLDER_7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model77. The consequences described are that later tokens are insufficiently influenced by the condition, text–image alignment deteriorates, and earlier errors are harder to correct.

The second pathology is over-guidance. When PRESERVED_PLACEHOLDER_7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model78 is too large, or when PRESERVED_PLACEHOLDER_7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model79 is large, the conditional branch dominates and can force semantics that conflict with the current visual context. The paper associates this with artifacts such as duplicated parts and unnatural shapes. Its illustrative example is a prompt emphasizing a “banana,” where strong guidance maps that attribute onto the elephant’s tusk, creating semantic misalignment. In LuminaGPT examples, standard CFG is described as producing tangled motorcycles, extra trunks, and redundant human hands, whereas SoftCFG is reported to reduce these artifacts. The paper explicitly likens the pair guidance diminishing versus over-guidance to gradient vanishing versus exploding in training, framing SoftCFG and Step Normalization as a regularization-and-normalization response to that duality (&&&7query7&&&).

7sort_by7. Uncertainty-guided perturbation of the value cache

SoftCFG’s core idea is to let each generated token contribute certainty-weighted guidance through the model’s internal memory. For each previously generated token PRESERVED_PLACEHOLDER_7max_results7query7, the method uses the maximum probability from the conditional branch as a confidence proxy:

PRESERVED_PLACEHOLDER_7max_results7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7^

and defines

PRESERVED_PLACEHOLDER_7max_results7max_results7^

Here PRESERVED_PLACEHOLDER_7max_results7sort_by7^ is described as an uncertainty weight. The paper states that tokens with higher confidence, corresponding to lower PRESERVED_PLACEHOLDER_7max_results7relevance7, are considered more reliable and thus receive stronger perturbations. The perturbation itself is applied to the unconditional value cache:

PRESERVED_PLACEHOLDER_7max_results7query7^

Using the perturbed unconditional cache PRESERVED_PLACEHOLDER_7max_results7\7, the model computes

PRESERVED_PLACEHOLDER_7max_results77^

SoftCFG then forms guided logits as

PRESERVED_PLACEHOLDER_7max_results78

which the paper also rewrites as

PRESERVED_PLACEHOLDER_7max_results79

In Algorithm 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7, the algebraically equivalent form is

PRESERVED_PLACEHOLDER_7sort_by7query7^

The term PRESERVED_PLACEHOLDER_7sort_by7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7^ is characterized as a context-aware regularizer. When it is aligned with the standard text-only guidance direction PRESERVED_PLACEHOLDER_7sort_by7max_results7, it amplifies effective guidance; when it is misaligned, it reduces the effective guidance magnitude and therefore counteracts over-guidance. Mechanistically, SoftCFG differs from standard CFG by touching every value-cache entry rather than only exploiting the conditional token at the beginning of the sequence. Because the value cache is reused at subsequent decoding steps through self-attention, the effect of confident tokens persists over time, embedding guidance into the AR model’s context memory rather than injecting it only as a per-step global logit offset (&&&7query7&&&).

The presentation includes a technical subtlety. Although PRESERVED_PLACEHOLDER_7sort_by7sort_by7^ is termed an uncertainty weight and the cache scaling uses PRESERVED_PLACEHOLDER_7sort_by7relevance7, the text also emphasizes that high-confidence tokens are the reliable guidance carriers. The paper’s algorithmic resolution is to use PRESERVED_PLACEHOLDER_7sort_by7query7^ as the quantity normalized by Step Normalization, so the normalized perturbation budget is distributed according to confidence rather than uncertainty.

7relevance7. Step Normalization and bounded long-sequence perturbation

SoftCFG alone introduces a cumulative-perturbation issue. Repeatedly scaling unconditional value vectors across a long sequence suppresses information in PRESERVED_PLACEHOLDER_7sort_by7\7; for high-confidence tokens, where PRESERVED_PLACEHOLDER_7sort_by77, the suppression becomes strong. The paper formalizes the resulting deviation between SoftCFG and vanilla CFG as

PRESERVED_PLACEHOLDER_7sort_by78

Assuming PRESERVED_PLACEHOLDER_7sort_by79 is PRESERVED_PLACEHOLDER_7relevance7query7-Lipschitz with respect to the value cache at step PRESERVED_PLACEHOLDER_7relevance7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7, the appendix gives the bound

PRESERVED_PLACEHOLDER_7relevance7max_results7^

Because PRESERVED_PLACEHOLDER_7relevance7sort_by7^ can grow with sequence length, the deviation may explode as decoding proceeds. The paper associates this with instability, drift or collapse of unconditional logits, and poor sample diversity. In its entropy analysis, unnormalized SoftCFG shows exploding normalized entropy at later steps (&&&7query7&&&).

Step Normalization is introduced to remove the sequence-length dependence by normalizing the effective confidence contributions at each step:

PRESERVED_PLACEHOLDER_7relevance7relevance7^

Algorithm 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7^ uses a small PRESERVED_PLACEHOLDER_7relevance7query7^ for numerical stability:

PRESERVED_PLACEHOLDER_7relevance7\7^

Defining PRESERVED_PLACEHOLDER_7relevance77, Step Normalization yields

PRESERVED_PLACEHOLDER_7relevance78

The perturbed values then become

PRESERVED_PLACEHOLDER_7relevance79

The paper’s proposition states that under the same PRESERVED_PLACEHOLDER_7query7query7-Lipschitz assumption,

PRESERVED_PLACEHOLDER_7query7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7^

This replaces a bound that grows with PRESERVED_PLACEHOLDER_7query7max_results7^ by one that depends only on the maximum value-vector norm. In the paper’s interpretation, the total perturbation budget is fixed to one unit at every step, so adding more tokens redistributes the budget rather than increasing it. This is the principal stabilization mechanism for long sequences and is reported to keep normalized entropy stable over generation steps (&&&7query7&&&).

7query7. Sampling procedure and implementation characteristics

SoftCFG assumes a trained AR image model with a VQ tokenizer, a decoder-only transformer with key/value caches, and standard CFG support. It operates with conditional input PRESERVED_PLACEHOLDER_7query7sort_by7, standard sampling hyperparameters such as temperature and top-PRESERVED_PLACEHOLDER_7query7relevance7/top-PRESERVED_PLACEHOLDER_7query7query7 guidance scale PRESERVED_PLACEHOLDER_7query7\7, and sequence length PRESERVED_PLACEHOLDER_7query77. The sampling loop described in Algorithm 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7^ is:

7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7. Initialize PRESERVED_PLACEHOLDER_7query78, empty conditional and unconditional caches, and a buffer for per-token confidence. 7max_results7. At each step PRESERVED_PLACEHOLDER_7query79, compute conditional logits PRESERVED_PLACEHOLDER_7\7query7. 7sort_by7. For each previous token PRESERVED_PLACEHOLDER_7\7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7, retrieve PRESERVED_PLACEHOLDER_7\7max_results7^ and form PRESERVED_PLACEHOLDER_7\7sort_by7. 7relevance7. Apply Step Normalization to obtain PRESERVED_PLACEHOLDER_7\7relevance7. 7query7. Rescale the unconditional value cache entrywise by PRESERVED_PLACEHOLDER_7\7query7. 7\7. Compute perturbed unconditional logits PRESERVED_PLACEHOLDER_7\7\7.

  1. Form PRESERVED_PLACEHOLDER_7\77, sample PRESERVED_PLACEHOLDER_7\78, and update both original caches and the stored max probability for the new token.
  2. After PRESERVED_PLACEHOLDER_7\79 steps, decode the token sequence to pixels with the VQ decoder.

The method is characterized as training-free and model-agnostic because it introduces no architectural changes, no new layers, and no tokenizer or decoder modifications. It requires access only to logits, the value cache, and the conditioning mechanism already used for CFG. Its added cost is described as per-step ztCFG=ztuncond+γ(ztcond−ztuncond),\mathbf{z}_{t}^{\text{CFG}} = \mathbf{z}_{t}^{\text{uncond}} + \gamma\big(\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\big),7query7^ scalar operations on the value cache, which are negligible relative to attention and MLP computation. The paper further reports that inference speed is almost identical to standard CFG on AliTok-7query7.7\7 and LuminaGPT7max_results7-7B (&&&7query7&&&).

A practical detail emphasized in the hyperparameter discussion is that SoftCFG can be used with scheduled guidance. The paper gives

ztCFG=ztuncond+γ(ztcond−ztuncond),\mathbf{z}_{t}^{\text{CFG}} = \mathbf{z}_{t}^{\text{uncond}} + \gamma\big(\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\big),7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7^

where ztCFG=ztuncond+γ(ztcond−ztuncond),\mathbf{z}_{t}^{\text{CFG}} = \mathbf{z}_{t}^{\text{uncond}} + \gamma\big(\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\big),7max_results7^ controls how guidance is distributed over the sequence. On AliTok-XL, the reported optimal ztCFG=ztuncond+γ(ztcond−ztuncond),\mathbf{z}_{t}^{\text{CFG}} = \mathbf{z}_{t}^{\text{uncond}} + \gamma\big(\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\big),7sort_by7^ is around ztCFG=ztuncond+γ(ztcond−ztuncond),\mathbf{z}_{t}^{\text{CFG}} = \mathbf{z}_{t}^{\text{uncond}} + \gamma\big(\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\big),7relevance7–ztCFG=ztuncond+γ(ztcond−ztuncond),\mathbf{z}_{t}^{\text{CFG}} = \mathbf{z}_{t}^{\text{uncond}} + \gamma\big(\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\big),7query7 The paper also notes that because SoftCFG “naturally amplifies” guidance via context perturbation, effective ztCFG=ztuncond+γ(ztcond−ztuncond),\mathbf{z}_{t}^{\text{CFG}} = \mathbf{z}_{t}^{\text{uncond}} + \gamma\big(\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\big),7\7^ values can be smaller than for vanilla CFG.

7\7. Empirical evaluation and qualitative behavior

The paper evaluates SoftCFG on class-conditional ImageNet-7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7K at ztCFG=ztuncond+γ(ztcond−ztuncond),\mathbf{z}_{t}^{\text{CFG}} = \mathbf{z}_{t}^{\text{uncond}} + \gamma\big(\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\big),7, where 7query7query7k images are sampled on the validation label set, and on text-to-image benchmarks GenEval and DPG-Bench. The class-conditional experiments use AliTok (B, L, XL), described in the paper as a state-of-the-art AR model on ImageNet ztCFG=ztuncond+γ(ztcond−ztuncond),\mathbf{z}_{t}^{\text{CFG}} = \mathbf{z}_{t}^{\text{uncond}} + \gamma\big(\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\big),8, and the text-to-image experiments use LuminaGPT7max_results7, described as a large AR image model with top performance on GenEval and DPG-Bench. Reported evaluation metrics for ImageNet include FID, Inception Score (IS), precision/recall, and sFID (&&&7query7&&&).

On the main ImageNet ztCFG=ztuncond+γ(ztcond−ztuncond),\mathbf{z}_{t}^{\text{CFG}} = \mathbf{z}_{t}^{\text{uncond}} + \gamma\big(\mathbf{z}_{t}^{\text{cond}}-\mathbf{z}_{t}^{\text{uncond}}\big),9 comparison, the paper reports that AliTok-XL baseline achieves FID γ>0\gamma>07query7, IS γ>0\gamma>07(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7, sFID γ>0\gamma>07max_results7, and recall γ>0\gamma>07sort_by7, whereas AliTok-XL + SoftCFG achieves FID 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7.7max_results77^, IS γ>0\gamma>07relevance7, sFID γ>0\gamma>07query7, and recall γ>0\gamma>07\7. Among AR models, this is presented as the best FID, surpassing RAR-XL at γ>0\gamma>07, MAR-H at γ>0\gamma>08, VAR-d7sort_by7query7-re at γ>0\gamma>09, RandAR-XXL at pθCFG(⋅∣x<t,c)=softmax(ztCFG).p_\theta^{\text{CFG}}(\cdot \mid x_{<t}, c) = \mathrm{softmax}(\mathbf{z}_{t}^{\text{CFG}}).7query7, and LlamaGen-7sort_by7B at pθCFG(⋅∣x<t,c)=softmax(ztCFG).p_\theta^{\text{CFG}}(\cdot \mid x_{<t}, c) = \mathrm{softmax}(\mathbf{z}_{t}^{\text{CFG}}).7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7. Relative to diffusion transformers, the paper lists DiT-XL/7max_results7^ at pθCFG(⋅∣x<t,c)=softmax(ztCFG).p_\theta^{\text{CFG}}(\cdot \mid x_{<t}, c) = \mathrm{softmax}(\mathbf{z}_{t}^{\text{CFG}}).7max_results7, MDTV7max_results7-XL/7max_results7^ at pθCFG(⋅∣x<t,c)=softmax(ztCFG).p_\theta^{\text{CFG}}(\cdot \mid x_{<t}, c) = \mathrm{softmax}(\mathbf{z}_{t}^{\text{CFG}}).7sort_by7, and REPA-E at pθCFG(⋅∣x<t,c)=softmax(ztCFG).p_\theta^{\text{CFG}}(\cdot \mid x_{<t}, c) = \mathrm{softmax}(\mathbf{z}_{t}^{\text{CFG}}).7relevance7, and describes AliTok-XL + SoftCFG’s FID of pθCFG(⋅∣x<t,c)=softmax(ztCFG).p_\theta^{\text{CFG}}(\cdot \mid x_{<t}, c) = \mathrm{softmax}(\mathbf{z}_{t}^{\text{CFG}}).7query7^ as statistically on par with state-of-the-art diffusion models.

The ablation study on AliTok-XL separates the contribution of SoftCFG and Step Normalization:

Setting FID Other reported metrics
Baseline (no CFG) 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7.77\7^ IS 7max_results7max_results7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7.7max_results7, sFID 7query7.7query7query7
+ CFG 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7.7sort_by77^ IS 7sort_by7max_results7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7.7relevance7, sFID 7.7max_results79
+ SoftCFG (no StepNorm) 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7.7sort_by7max_results7^ IS 7max_results788.7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7, sFID 7.7\7max_results7^
+ SoftCFG + StepNorm 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7.7sort_by7max_results7^ IS 7sort_by7query7max_results7.7query7 sFID 7.7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7\7^
+ SoftCFG + StepNorm + Opt. 7(Xu et al., 1 Oct 2025) SoftCFG uncertainty-guided stable guidance for visual autoregressive model7.7max_results77^ IS 7sort_by7query7max_results7.7relevance7 sFID 7\7.77query7

The stated interpretation is that SoftCFG improves FID over vanilla CFG even without Step Normalization, while Step Normalization stabilizes IS and improves sFID. Joint tuning of pθCFG(⋅∣x<t,c)=softmax(ztCFG).p_\theta^{\text{CFG}}(\cdot \mid x_{<t}, c) = \mathrm{softmax}(\mathbf{z}_{t}^{\text{CFG}}).7\7^ and the cosine exponent pθCFG(⋅∣x<t,c)=softmax(ztCFG).p_\theta^{\text{CFG}}(\cdot \mid x_{<t}, c) = \mathrm{softmax}(\mathbf{z}_{t}^{\text{CFG}}).7 yields the best results.

Qualitatively, the paper attributes to SoftCFG better structural consistency and fewer artifacts under strong guidance. In LuminaGPT examples, SoftCFG is described as reducing tangled motorcycles, extra trunks, and extra hands. Confidence heatmaps align high-confidence regions with salient semantic parts such as faces and object cores, while ambiguous backgrounds receive lower confidence. The paper treats this as evidence that high-confidence tokens are effective guidance carriers. At the same time, the method exhibits a documented failure mode: for the prompt describing “A calico cat… patchwork of orange, black, and white fur… on top of a sleek Mercedes Benz…”, the patchwork pattern is applied to the car rather than the cat. The accompanying confidence map reportedly places high confidence on the car and low confidence on the cat, illustrating that SoftCFG depends on the quality and localization of the model’s confidence estimates (&&&7query7&&&).

7. Relation to adjacent guidance methods, limitations, and prospective extensions

Within the paper’s taxonomy, SoftCFG is positioned as an AR analogue of uncertainty- or semantics-aware guidance methods that modulate guidance internally rather than relying solely on a fixed global logit offset. The paper links it conceptually to semantic-aware CFG, token-level guidance for diffusion, self-attention guidance, perturbed-attention guidance, and smoothed energy guidance. Its stated novelty is that this philosophy is adapted to AR generation by operating on the value cache, which serves as the primary context memory in decoder-only AR models. The paper also contrasts SoftCFG with CCA, which seeks guidance-free AR sampling through retraining, and with architectural families such as RAR, VAR, and MAR, which explore AR conditioning strategies but do not focus on uncertainty-aware inference (&&&7query7&&&).

The limitations identified in the paper are specific. First, the method relies on confidence quality: it assumes that larger pθCFG(⋅∣x<t,c)=softmax(ztCFG).p_\theta^{\text{CFG}}(\cdot \mid x_{<t}, c) = \mathrm{softmax}(\mathbf{z}_{t}^{\text{CFG}}).8 corresponds to better semantic consistency, yet AR confidence estimates are not always calibrated or well localized. The calico-cat failure case is presented as direct evidence. Second, Step Normalization may be rigid: because it enforces pθCFG(⋅∣x<t,c)=softmax(ztCFG).p_\theta^{\text{CFG}}(\cdot \mid x_{<t}, c) = \mathrm{softmax}(\mathbf{z}_{t}^{\text{CFG}}).9, the perturbation is capped to the equivalent of at most one token’s worth per step, which may under-utilize guidance in very long sequences. Third, although Step Normalization removes linear growth in the deviation bound, extremely long horizons such as ultra-high resolution generation or video may require more adaptive normalization. Fourth, the experimental scope is limited to class-conditional and text-to-image image generation; multimodal AR tasks, video synthesis, and RLHF-trained generators remain outside the reported evaluation.

The paper’s suggested future directions follow from those limitations. It proposes using external perceptual models such as DINOv7sort_by7^ for better semantic confidence estimation, exploring adaptive or hierarchical normalization that scales the perturbation budget with context complexity or length, and extending the underlying principle—that previously generated visual content can guide subsequent tokens—to other generative systems, including LLMs for text and AR video models. This suggests that SoftCFG is best understood not only as a particular AR image-sampling heuristic, but as a broader proposal for integrating internally derived confidence signals into guidance mechanisms while preserving the training-free structure of CFG.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SoftCFG.