Energy-Based Decoding
- Energy-based decoding is a family of techniques that integrates energy functions into the candidate generation and full-sequence inference process.
- In neural text generation, methods like COLD and EBD replace traditional left-to-right token selection with global optimization using Langevin dynamics and Metropolis–Hastings sampling.
- Applications span LVLM layer selection, video codec rate-energy-distortion optimization, and physical energy measurement in hardware, highlighting practical trade-offs in performance and efficiency.
Searching arXiv for recent and foundational papers on energy-based decoding across text generation, LVLMs, and codec/system interpretations. arXiv_search.query({"search_query":"all:\"Energy-based Decoding\" OR ti:\"COLD Decoding\" OR ti:\"Energy-Guided Decoding\" OR ti:\"Decoding-Energy-Rate-Distortion\"","max_results":10,"sort_by":"relevance","sort_order":"descending"}) Energy-based decoding is a family of decoding-time methods in which an energy quantity influences candidate generation, selection, or refinement. In contemporary research, the term is used in several technically distinct ways. In neural text generation, it often denotes a scalar defined over complete sequences, so decoding becomes approximate inference in an energy-based model rather than left-to-right token choice; representative formulations include and (Qin et al., 2022, Wang et al., 27 May 2026). In large vision-LLMs, it can denote a per-layer logit score $Energy(h_t^k)=-\logsumexp[\mathcal H(h_t^k)]$ used to choose the layer from which the next token is decoded (Liu et al., 10 Jul 2025). In video coding and systems work, the same phrase instead refers to minimizing the physical energy required by the decoder, for example through objectives such as or by comparing the GPU energy of alternative decoding policies (Herglotz et al., 2022, Nik et al., 17 Feb 2025). In communication-theoretic and VLSI analyses, it refers even more literally to the physical energy of iterative decoders and decoder circuits (Rehman et al., 2013, Blake et al., 2014).
1. Terminological scope and recurrent structures
Across the cited literature, “energy-based decoding” does not identify a single formalism. It names several practices in which a scalar energy, or a measured/estimated physical energy, becomes part of the decoding rule. The commonality is not the mathematical form alone, but the relocation of decoding away from pure local likelihood maximization toward a broader optimization or inference problem.
| Usage in the literature | Energy quantity | Representative papers |
|---|---|---|
| Sequence-level neural decoding | Global sequence energy or reward-tilted energy | (Qin et al., 2022, Wang et al., 27 May 2026) |
| LVLM layer selection | Per-layer logit energy $-\logsumexp$ | (Liu et al., 10 Jul 2025) |
| Video codec optimization | Predicted decoder processing energy | (Herglotz et al., 2022, Kränzler et al., 2022, Herglotz et al., 2022, Kränzler et al., 2022) |
| Systems and hardware studies | Measured GPU energy, iterative decoding energy, or VLSI area-time energy | (Nik et al., 17 Feb 2025, Rehman et al., 2013, Blake et al., 2014) |
This terminological breadth matters because superficially similar language can refer to very different decoded objects. In COLD and EBD, the decoded object is a full text response or a relaxed soft-token sequence; in the LVLM setting, the decoded object remains the next token, but the decoding layer changes dynamically; in codec papers, the immediate decision variable is often an encoder-side mode or tool choice chosen to reduce downstream decoder energy; and in hardware theory, the object of study is the physical energy required to perform decoding at all (Qin et al., 2022, Herglotz et al., 2022, Blake et al., 2014).
2. Sequence-level energy models in text generation
A canonical sequence-level formulation appears in COLD decoding, which casts constrained text generation as inference in an energy-based model over whole sequences rather than as left-to-right token selection. Constraint scores are combined into
with nonnegative weights . Because text is discrete, COLD relaxes each token position into a logit vector , forms soft tokens with temperature-scaled softmax, and performs Langevin dynamics over the full relaxed sequence: The framework uses a left-to-right LM fluency term, can optionally add a reverse-LM term, and instantiates additional differentiable energies such as future-token prediction and differentiable 0-gram similarity. It is applied without task-specific fine-tuning to lexically constrained generation, abductive reasoning, and counterfactual reasoning. The reported procedure initializes from greedy LM decoding logits, typically uses 1 and 2, anneals 3 through 4 at iterations 5, and finally discretizes with an LM-guided top-6 “guardian” rule rather than naive argmax. Empirically, the method outperforms DeLorean and left-only LM baselines on abductive reasoning, improves coherence over DeLorean on counterfactual rewriting, significantly outperforms Mix-and-Match, and achieves higher keyword coverage than NeuroLogic, while exhibiting a recurring fluency–constraint tradeoff; runtime is reported as about 7 seconds per sample with GPT2-XL and 8 seconds with GPT2-M on counterfactual rewriting (Qin et al., 2022).
A later sequence-level variant, explicitly titled Energy-Based Decoding, replaces soft-token Langevin refinement with reward-guided posterior sampling over complete responses from frozen base LLMs. EBD begins from the KL-regularized objective
9
whose optimizer is the reward-tilted target
$Energy(h_t^k)=-\logsumexp[\mathcal H(h_t^k)]$0
In energy form,
$Energy(h_t^k)=-\logsumexp[\mathcal H(h_t^k)]$1
The implementation uses a prompt-normalized standardized advantage
$Energy(h_t^k)=-\logsumexp[\mathcal H(h_t^k)]$2
initializes from a small pool of prior samples, and refines responses with a short block-wise Metropolis–Hastings chain that preserves a prefix, regenerates a suffix from the matched conditional prior, and accepts proposals with
$Energy(h_t^k)=-\logsumexp[\mathcal H(h_t^k)]$3
Because the proposal matches the conditional prior, proposal and prior terms cancel in the MH ratio. Default settings are $Energy(h_t^k)=-\logsumexp[\mathcal H(h_t^k)]$4, $Energy(h_t^k)=-\logsumexp[\mathcal H(h_t^k)]$5, $Energy(h_t^k)=-\logsumexp[\mathcal H(h_t^k)]$6, $Energy(h_t^k)=-\logsumexp[\mathcal H(h_t^k)]$7, $Energy(h_t^k)=-\logsumexp[\mathcal H(h_t^k)]$8, and $Energy(h_t^k)=-\logsumexp[\mathcal H(h_t^k)]$9. On five base models and six benchmarks, EBD improves both objective and subjective evaluations relative to direct decoding and Power Sampling; the abstract highlights Qwen3-8B-Base on AlpacaEval2.0 from 0 to 1, an 2 Math500 latency reduction for Mistral-7B relative to prior decoding work, and robustness to reward-model size from 3B to 4B (Wang et al., 27 May 2026).
Taken together, these two frameworks show two distinct sequence-level regimes. COLD performs approximate sampling in a continuous relaxation of text with Langevin dynamics. EBD keeps text discrete and instead samples from a reward-tilted posterior with matched block-wise MH proposals. Both reject the assumption that decoding should be identical to ancestral next-token selection, but they do so with markedly different inference mechanisms (Qin et al., 2022, Wang et al., 27 May 2026).
3. Energy-guided layer selection in vision-language decoding
In large vision-LLMs, energy-guided decoding is neither an EBM over full responses nor an optimization over physical power. The method introduced for object hallucination mitigation inspects hidden states from all decoder layers at each generation step, projects each layer’s last-position hidden state through the existing LM head, and computes the energy score
5
The chosen layer is
6
and decoding then proceeds from 7. The method is hyperparameter-free in the energy definition, requires no retraining, no finetuning, no extra model, no visual perturbation, no prompt tuning, and no contrastive pair of distributions. It is evaluated on LLaVA-1.5, InstructBLIP, and mPLUG-Owl2 across POPE, MME, MMVP, and CHAIR, with a maximum of 16 new tokens for yes/no evaluations and temperature 8 (Liu et al., 10 Jul 2025).
The motivating observation is a strong “Yes”-ratio imbalance in balanced yes/no VQA datasets. The paper argues that ordinary final-layer decoding often over-predicts “Yes,” partly due to language-prior transfer, and that the final layer is not always the most reliable source for the next-token distribution. Selecting the minimum-energy layer is presented as a way to reduce that bias and improve calibration. Empirically, the abstract reports an average accuracy improvement of 9 over greedy decoding and an average yes-ratio gap reduction of $-\logsumexp$0. The detailed tables show especially large gains on harder POPE settings such as LLaVA-1.5 on GQA adversarial, where accuracy rises from $-\logsumexp$1 to $-\logsumexp$2, F1 from $-\logsumexp$3 to $-\logsumexp$4, and $-\logsumexp$5 drops from $-\logsumexp$6 to $-\logsumexp$7. On MME, LLaVA-1.5 improves from $-\logsumexp$8 to $-\logsumexp$9, and on MMVP the method substantially reduces yes-ratio bias, although F1 changes are sometimes small. The layer-wise analysis indicates that the selected hidden states mostly come from the second-last layer, and the penultimate layer often has the lowest energy (Liu et al., 10 Jul 2025).
The method is not uniformly dominant in every condition. Some easier MSCOCO POPE settings do not improve on every metric, and in open-ended captioning the method gives the best CHAIR0 but worse CHAIR1; for LLaVA-1.5, CHAIR2 changes from 3 to 4 while CHAIR5 changes from 6 to 7. This establishes a narrower but precise meaning of energy-guided decoding: a layer-selection rule defined by log-sum-exp over vocabulary logits, used chiefly as a reliability and calibration signal during autoregressive decoding (Liu et al., 10 Jul 2025).
4. Decoder-energy-rate-distortion optimization in video coding
In video coding, energy-based decoding refers primarily to decoder-energy-aware encoding. The foundational DERDO formulation extends classical rate-distortion optimization by treating decoder processing energy as a first-class objective. Over feasible encoding solutions 8, the constrained problem is
9
with Lagrangian relaxation
0
A central practical contribution is the encoder-side parameter 1, which continuously shifts operation from pure RDO 2 toward pure decoding-energy optimization 3. The paper integrates this objective into a modified HM-14.0 HEVC encoder using a 27-feature bitstream model trained on FFmpeg measurements, and validates the two-multiplier QP relation
4
Measured on a Pandaboard with FFmpeg 2.8, libde265 0.7, and HM-13.0, the method yields up to about 5 measured decoder energy reduction at equal objective quality, with bitrate increase of similar order; the paper reports, for example, class-averaged FFmpeg results at 6 of about 7–8 BDDE with only 9–0 BDR increase, and at 1 about 2–3 BDDE with 4–5 BDR penalties (Herglotz et al., 2022).
Such optimization depends on accurate energy estimation. A closely related line of work models video decoding energy directly from bitstream features using
6
For HEVC, H.264, H.263, and VP9, fewer than 20 features suffice: 7 for H.263, 8 for H.264, and 9 for HEVC and VP9. Using 10-fold cross-validation and external measurements with a ZES Zimmer LMG95 on a Pandaboard, mean relative estimation errors on FFmpeg are 0 for HEVC, 1 for H.264, 2 for H.263, and 3 for VP9; on alternative implementations the method reaches 4 for HEVC on libde265, 5 for H.263 on TMN-2.0, and 6 for a hardware-accelerated H.264 decoder. The paper uses these results to support rate-distortion-energy analysis before playback rather than to define a decoder-side probabilistic energy (Herglotz et al., 2022).
For VVC, the same modeling philosophy is extended with richer feature sets. The paper on VVC decoding-energy modeling emphasizes that VVC random-access decoding energy increases by over 7 on average relative to HEVC and proposes two feature-based models, FVS with 67 features and FV with 230 features, both learned by trust-region-reflective least squares from VTM-7.0 measurements on an Intel i7-8700 using RAPL. With 10-fold cross-validation, FV achieves 8 mean relative estimation error on the merged dataset, which the paper frames as sufficiently accurate for subsequent DERDO-style optimization (Kränzler et al., 2022).
A more deployment-oriented VVC study moves from modeling to practical coding-tool profiles in VVenC 1.3.1 and VVdeC 1.3.0. Starting from an aggressive Energy Efficient profile and then selectively restoring tools such as restricted ALF/CCALF, deeper partitioning, affine/PROF, and LMChroma, it defines two recommended profiles. The bitrate-efficient profile v2568 achieves JVET CTC averages of 9 and 0, matching the abstract’s statement of over 1 energy-efficiency improvement with bitrate increase below 2. The more aggressive profile v58 achieves 3 and 4, corresponding to 5 energy savings at bitrate increases below 6. Here again, the “energy-based” component is chiefly encoder-side control of decoder-expensive syntax and tools rather than a decoder-side inference rule (Kränzler et al., 2022).
5. Physical energy of decoding strategies and decoder circuits
A systems-oriented interpretation appears in work on GPU energy consumption during LLM inference. This literature is explicit that it is not about energy-based models, but about the physical energy implications of decoding choices. A comparative study on Qwen2.5-7B-Instruct evaluates 12 strategies—greedy, beam search, diverse beam search, contrastive search, DoLa, assisted decoding, temperature sampling, top-7, top-8, epsilon, typical, and min-9—over WMT16 translation, CodeXGLUE code summarization, and GSM8K, measuring GPU energy with 00 from 01 s nvidia-smi samples. It reports that decoding choice often changes energy much more than it changes task quality, that beam and DBS typically improve metrics at substantial energy cost, that contrastive search can be quality-optimal but very energy-expensive, and that assisted decoding often yields the best efficiency ratio 02. On German03English, for example, beam search reaches BLEU 04 at 05 Wh, while assisted decoding reaches BLEU 06 at 07 Wh and the best ER 08; on GSM8K, beam reaches 09 exact match at 10 Wh, whereas assisted decoding reaches 11 at 12 Wh and the best ER 13 (Nik et al., 17 Feb 2025).
Communication and coding theory use the term even more literally. In Wireless Body Area Sensor Networks, iterative LDPC decoding is analyzed from the viewpoint that decoder energy can exceed transmit energy on short links. The paper models decoder energy as increasing linearly with the number of computational nodes and iterations, lower bounded by
14
and proposes Adaptive Iterative Decoding, an early-stopping rule that halts once a BER target of 15 is reached. At 16 dB, the required iteration counts are 17 for 18 and 19 for 20; compared with a 50-iteration baseline, the paper reports total energy reduction of 21–22 and gives an example in which energy drops from about 23 J to about 24 J (Rehman et al., 2013).
A different communication-theoretic usage appears in bit-interleaved coded energy-based modulation with iterative decoding. There, “energy-based” refers to modulation symbols whose information is carried in energy levels under non-coherent reception, and the decoding problem is solved with an iterative BICM-ID-style receiver. The central non-coherent likelihood is
25
and the paper derives FF and EFF pairwise-error bounds, shows that nearest neighbors differ from the coherent BICM-ID case, and proves that the mapping from bits to energy levels influences diversity order and coding gain (Fazeli et al., 2022).
At the most abstract end, VLSI theory studies the physical energy of decoder circuits through Thompson’s model, where
26
For families of circuits decoding over a binary erasure channel, the paper proves that, as blocklength grows, either block error probability becomes asymptotically lower bounded by 27 or total decoding energy scales at least as 28, implying energy per decoded bit 29. For serial computation with a constant number of output pins, the lower bound strengthens to 30; in a more general setting with varying output pins, it becomes 31. A further implication is that the average energy per decoded bit must approach infinity for any sequence of codes that approaches capacity (Blake et al., 2014).
6. Trade-offs, misconceptions, and comparative interpretation
The surveyed literature establishes that “energy” in decoding is not a single object. It may be a global score over whole text sequences, a reward-tilted posterior energy, a per-layer logit reliability score, a predicted decoder processing cost extracted from bitstream features, the actual GPU energy used by an inference policy, the iteration-dependent energy of a message-passing decoder, or the area-time energy of a VLSI circuit (Qin et al., 2022, Liu et al., 10 Jul 2025, Herglotz et al., 2022, Nik et al., 17 Feb 2025, Blake et al., 2014). A persistent misconception is therefore to equate all energy-based decoding with energy-based models; the codec and systems literature uses the same phrase for physically measured or estimated energy without introducing an EBM over outputs.
A second recurring pattern is that energy-aware decoding nearly always exposes a trade-off rather than a uniformly dominant operating point. In COLD, stronger task constraints can lower grammaticality, and top-32 discretization mediates grammar against right/overall coherence (Qin et al., 2022). In reward-guided EBD, the prior–reward balance is controlled by 33, and the method remains only an approximate sampler because it runs a short MH chain (Wang et al., 27 May 2026). In LVLMs, energy-guided layer selection reduces yes-ratio bias and improves calibration, but CHAIR34 can worsen even when CHAIR35 improves (Liu et al., 10 Jul 2025). In HEVC and VVC, lower decoder energy is exchanged against bitrate and sometimes encoding time (Herglotz et al., 2022, Kränzler et al., 2022). In energy-conscious LLM serving, beam-like methods often buy modest quality gains at disproportionate energy cost, while assisted decoding and simple stochastic truncation rules offer better quality-per-energy trade-offs (Nik et al., 17 Feb 2025). In short-range communication, aggressive transmit-power minimization can backfire because it raises decoding iterations and hence decoding energy (Rehman et al., 2013). In circuit theory, approaching capacity forces diverging energy per decoded bit (Blake et al., 2014).
A plausible synthesis is that the literature has converged on a shared design principle rather than a single formalism: decoding should be treated as an optimization or inference stage whose objective need not be identical to local next-token likelihood or local rate-distortion cost. What differs across domains is the status of the energy term—probabilistic, heuristic, predictive, or physical. That distinction determines both the mathematical tools used, such as Langevin dynamics, Metropolis–Hastings, log-sum-exp layer scoring, or Lagrangian rate-energy-distortion optimization, and the failure modes that dominate in practice, such as discretization error, reward dependence, calibration drift, bitrate penalty, or hardware energy overhead (Qin et al., 2022, Wang et al., 27 May 2026, Liu et al., 10 Jul 2025, Herglotz et al., 2022, Nik et al., 17 Feb 2025).