Papers
Topics
Authors
Recent
Search
2000 character limit reached

Finite-Sample Redundancy Laws in Information Theory

Updated 26 September 2025
  • Finite-sample redundancy laws are quantitative relationships that characterize the excess inefficiency in algorithms due to limitations like finite precision, delay, and sample size.
  • They reveal how redundancy scales with specific parameters in contexts such as source coding, universal compression, and deep learning.
  • These laws guide practical design choices by balancing trade-offs in resource allocation, system robustness, and performance optimization across diverse applications.

Finite-sample redundancy laws refer to rigorous quantitative relationships that characterize the excess penalty—in terms of expected code length, risk, or representational inefficiency—incurred in various algorithms, models, and physical systems due to non-asymptotic, resource-constrained, or "imperfect" conditions such as finite precision, finite delay, finite sample size, finite blocklength, or structural limitations. These laws establish how redundancy scales with problem parameters, algorithmic choices, and implementation constraints, and are foundational for understanding optimality, robustness, and resource allocation in information theory, coding, compression, statistical inference, signal processing, and learning theory.

1. Precision–Redundancy Tradeoffs in Source Coding

Finite-precision representation of source probabilities directly leads to excess redundancy in classic source coding algorithms such as Shannon, Gilbert-Moore, Huffman, and arithmetic codes. For a source with alphabet size mm, probabilities pip_i are approximated by rationals fi/tf_i/t stored with WW bits, yielding a redundancy RR that satisfies the subadditive bound: Wηlog2mR,W \lesssim \eta \log_2 \frac{m}{R}, where η\eta is an implementation-dependent constant (12\frac{1}{2} for binary sources, m/(m+1)m/(m+1) for optimized mm-ary codes, pip_i0 for general progressive update designs) (0712.0057). The Kullback–Leibler divergence pip_i1 is bounded via the maximal approximation error pip_i2 as pip_i3, translating the effect of denominator pip_i4 (and hence pip_i5) to the residual redundancy. The binary case admits Diophantine-optimal approximations with redundancy decaying as pip_i6 (leading to a halved pip_i7), while pip_i8-ary cases exhibit redundancy decay as pip_i9, with practical code design implications for memory, hardware register width, and symbol grouping.

2. Delay–Redundancy Laws in Lossless Source Coding

Imposing a finite decoding delay fi/tf_i/t0 on lossless source codes fundamentally affects the redundancy decay rate. In block/phrase-constrained coding (e.g., Huffman, Tunstall), redundancy decays polynomially with block/phrase length (O(1/d)); in contrast, delay-constrained sequential encoders (e.g., delay-limited arithmetic coding with bit flushing) achieve exponential decay: fi/tf_i/t1 where fi/tf_i/t2 is the Rényi entropy of order 2 of the source (Shayevitz et al., 2010). The redundancy-delay exponent fi/tf_i/t3, defined as fi/tf_i/t4, is lower-bounded by fi/tf_i/t5, but for almost all sources, it cannot exceed a bound depending on the minimal symbol probability and alphabet size. This exponential scaling marks a qualitative improvement over classical codes, and optimal code design under delay constraints is inextricably linked to the fine-grain properties of fi/tf_i/t6.

3. Redundancy Laws in Universal Data Compression on Countable Alphabets

For universal coding over a countably infinite alphabet, redundancy for a class fi/tf_i/t7 depends crucially on tail behavior. Finite single-letter redundancy (i.e., existence of fi/tf_i/t8 with fi/tf_i/t9) implies tightness, but not necessarily diminishing per-symbol redundancy with blocklength (Hosseini et al., 2014, Hosseini et al., 2018). The asymptotic per-symbol redundancy WW0 equals the tail redundancy: WW1 revealing that the cost of compressing novel, "tail" symbols dominates as WW2 grows: finite single-letter redundancy does not guarantee WW3, and only classes with vanishing tail redundancy are strongly compressible. This formalism captures the true essence of finite-sample redundancy in infinite-alphabet compression.

4. Minimax Redundancy and Regret in Parametric Models

In smooth parametric families (e.g., exponential families), finite-sample minimax redundancy and regret are determined by the Shtarkov and Jeffreys integrals (0903.5399, Beirami et al., 2011). For a WW4-parameter family, the worst-case redundancy exhibits the canonical scaling: WW5 where WW6 is the Jeffreys correction term. Sufficient conditions for finite redundancy include restriction to compact parameter sets and tail decay of the base measure (density WW7 for some WW8). For universal codes (including two-stage codes), the asymptotic average minimax redundancy serves as an accurate benchmark, while additional penalty terms for two-stage coding become negligible for large WW9. In nonstandard settings (e.g., mixtures with heavy tails), the Jeffreys integral may diverge, limiting applicability of classic finite-sample redundancy laws.

5. Pseudocodeword Redundancy in Linear Codes

Pseudocodeword redundancy measures the minimum number of parity-check rows in a matrix RR0 so that all non-zero pseudocodewords have weight at least RR1, the code’s minimum Hamming distance (Zumbragel et al., 2010, Zumbrägel et al., 2011). For iterative or LP decoding, this represents the finite-sample constraint needed to eliminate low-weight pseudocodewords and match ML decoding performance. Most random codes exhibit infinite pseudocodeword redundancy, but for codes based on designs (e.g., BIBDs) and cyclic codes meeting the Vontobel–Koetter eigenvalue bound,

RR2

finite redundancy is attainable. This trade-off connects structural code properties and practical decoder design in finite regimes.

6. Redundancy Allocation Laws in Partitioned Codes

For finite-length nested (partitioned) codes in nonvolatile memory applications, redundancy must be allocated between defect masking (RR3 bits) and error correction (RR4 bits), under constraints RR5 (Kim et al., 2018). Recovery failure probability is bounded as: RR6 where RR7 is the defect probability, RR8 is the erasure or crossover probability. The optimal allocation is estimated analytically (by KKT conditions) and matches simulation optima, underscoring the non-triviality of finite-sample performance compared to asymptotic theory.

7. Redundancy Laws in Structural Optimization and Function Approximation

Structural redundancy, formalized in robust optimization via information-gap theory, quantifies the maximal degradation (RR9) sustainable without exceeding performance thresholds, with worst-case performance Wηlog2mR,W \lesssim \eta \log_2 \frac{m}{R},0 (Kanno, 2016). Multiple damage scenarios yield non-differentiable optimization landscapes; algorithmic approaches such as derivative-free SQP leverage finite-difference gradients to navigate these constraints efficiently. In linear function approximation with numerically redundant bases (e.g., frames or overcomplete dictionary), numerical regularization (e.g., Wηlog2mR,W \lesssim \eta \log_2 \frac{m}{R},1 or TSVD) reduces required sample size, replacing the nominal dimensionality Wηlog2mR,W \lesssim \eta \log_2 \frac{m}{R},2 with an effective dimension Wηlog2mR,W \lesssim \eta \log_2 \frac{m}{R},3 such that

Wηlog2mR,W \lesssim \eta \log_2 \frac{m}{R},4

for accurate recovery (Herremans et al., 13 Jan 2025).

8. Redundancy Laws in Function-Correcting Codes and Feature Learning

Function-correcting codes over finite fields require redundancy Wηlog2mR,W \lesssim \eta \log_2 \frac{m}{R},5 (Ly et al., 19 Apr 2025). In large fields (Wηlog2mR,W \lesssim \eta \log_2 \frac{m}{R},6), optimal systematic MDS codes achieve Wηlog2mR,W \lesssim \eta \log_2 \frac{m}{R},7, while in binary and moderate-sized fields,

Wηlog2mR,W \lesssim \eta \log_2 \frac{m}{R},8

demonstrating logarithmic overhead with dimension. These explicit finite-length laws guide practical code constructions.

In deep learning, finite-sample scaling laws are shown to be redundancy laws (Bi et al., 25 Sep 2025). Kernel regression under a covariance spectrum with Wηlog2mR,W \lesssim \eta \log_2 \frac{m}{R},9 yields excess risk decaying as η\eta0 where

η\eta1

with η\eta2 (source condition) and redundancy parameter η\eta3. Universality is established across invertible transforms, mixture domains, finite-width models, and Transformers, demonstrating that the scaling exponent is not universal but dictated by data redundancy.

Summary Table: Key Redundancy Laws and Scaling

Context Scaling Law / Bound Governing Parameters
Precision–Redundancy (Source coding) η\eta4 η\eta5: code-dependent constant, η\eta6, η\eta7
Delay–Redundancy (Sequential codes) η\eta8 η\eta9: Rényi entropy order 2, 12\frac{1}{2}0
Universal coding (infinite alphabet) 12\frac{1}{2}1 Tail redundancy 12\frac{1}{2}2
Minimax Redundancy in Parametric Models 12\frac{1}{2}3 12\frac{1}{2}4: param dim., 12\frac{1}{2}5: Jeffreys integral
Partitioned/Nested Codes 12\frac{1}{2}6 12\frac{1}{2}7, 12\frac{1}{2}8, 12\frac{1}{2}9, m/(m+1)m/(m+1)0
Function Approximation (Frames) m/(m+1)m/(m+1)1 m/(m+1)m/(m+1)2: eff. dim via regularization
Function-Correcting Codes m/(m+1)m/(m+1)3 m/(m+1)m/(m+1)4: dim., m/(m+1)m/(m+1)5: error level
Deep Learning Scaling (Redundancy Law) m/(m+1)m/(m+1)6 m/(m+1)m/(m+1)7: smoothness, m/(m+1)m/(m+1)8: spectral tail

Finite-sample redundancy laws reveal the precise mechanisms by which resource constraints and discrete, non-asymptotic phenomena induce excess risk, inefficiency, or code length, and provide critical guidance for algorithm and system design across multiple disciplines. These laws unify previously disparate observations on scaling, robustness, and regularization, making explicit the fundamental role of redundancy in practical applications.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Finite-Sample Redundancy Laws.