Papers
Topics
Authors
Recent
Search
2000 character limit reached

Generalized Forms of the Kraft Inequality for Finite-State Encoders

Published 23 Jan 2026 in cs.IT | (2601.16594v1)

Abstract: We derive a few extended versions of the Kraft inequality for information lossless finite-state encoders. The main basic contribution is in defining a notion of a Kraft matrix and in establishing the fact that a necessary condition for information losslessness of a finite-state encoder is that none of the eigenvalues of this matrix have modulus larger than unity, or equivalently, the generalized Kraft inequality asserts that the spectral radius of the Kraft matrix cannot exceed one. For the important special case where the FS encoder is irreducible, we derive several equivalent forms of this inequality, which are based on well known formulas for spectral radius. It also turns out that in the irreducible case, Kraft sums are bounded by a constant, independent of the block length, and thus cannot grow even in any subexponential rate. Finally, two extensions are outlined - one concerns the case of side information available to both encoder and decoder, and the other is for lossy compression.

Authors (1)

Summary

  • The paper presents a new spectral-radius condition substituting Ziv–Lempel's 1978 asymptotic super-alphabet-based generalized Kraft inequality, showing that information losslessness of a finite-state encoder is equivalent to the spectral radius of a Kraft matrix.
  • The key finding demonstrates that the spectral radius of the Kraft matrix is a necessary condition that is proven to be equivalent to the classical Kraft inequality when the number of states is one, thereby providing a unified framework for both lossless and lossy coding.
  • The paper shows that reducible encoders confer no asymptotic advantage, the Craft matrix preserves topological probabilistic eigenvalues bounded on reducible and irreducible graphs, implying a uniformly bounded $2^{(s-1)L_{ ext{max}}$ for $K^n$.

Motivation and relation to prior work

Kraft's inequality, extended by McMillan to uniquely decodable codes, is the classical necessary and sufficient condition on codeword lengths for variable-length lossless codes, and underpins the converse to the source coding theorem. When the encoder has memory — modeled as a finite-state (FS) encoder in the sense of Ziv and Lempel (2601.16594), whose output and next state are determined recursively from the current symbol and state via functions ff and gg — the scalar Kraft inequality no longer applies directly. The only prior generalization for information-lossless (IL) FS encoders appears as Lemma 2 of Ziv–Lempel's 1978 paper, which bounds the block-level Kraft sum over \ell-vectors by s2[1+log(1+α/s2)]s^2[1+\log(1+\alpha^\ell/s^2)]. The paper identifies two defects of this formulation: it does not recover the classical Kraft inequality when s=1s=1 (the right-hand side exceeds unity even for =1\ell=1), and it is posed only at the level of super-alphabet extensions rather than at the single-symbol/state level at which the encoder is defined.

The Kraft matrix and the spectral-radius condition

The central construction associates to any IL FS encoder E=(X,Y,Z,f,g)E=(\mathcal{X},\mathcal{Y},\mathcal{Z},f,g) an s×ss\times s nonnegative Kraft matrix KK, with entries

Kzz={x: g(z,x)=z}2L[f(z,x)].K_{zz'}=\sum_{\{x:\ g(z,x)=z'\}}2^{-L[f(z,x)]}.

The main theorem asserts that information losslessness implies gg0, i.e., no eigenvalue of gg1 may have modulus exceeding one. The proof shows that each entry of gg2 equals gg3, which the IL property (injectivity of gg4 given fixed endpoints) bounds linearly in gg5, namely by gg6. If instead gg7, Perron–Frobenius theory supplies a nonnegative right eigenvector gg8, domination of gg9 by a multiple of \ell0, and hence exponential growth of \ell1 — contradicting the linear bound. This condition reduces exactly to the classical Kraft inequality for \ell2, since \ell3 degenerates to the ordinary Kraft sum.

An illustrative three-state implementation of the length-2 block code \ell4 has eigenvalues \ell5: the row sum corresponding to the start state exceeds unity, demonstrating that a general IL FS encoder need not satisfy per-state prefix-code Kraft conditions — the FS model is strictly more general than a machine implementing a separate UD code at every state. Whether \ell6 is also sufficient for information losslessness is left open; sufficiency holds trivially if every state satisfies its own Kraft inequality, and at the block level one can always prepend \ell7 bits to obtain a genuine prefix code.

Irreducible encoders: bounded Kraft sums

For irreducible encoders — those whose adjacency graph is strongly connected — the paper derives several equivalent Collatz–Wielandt forms of the GKI, including the statement that for every \ell8 there exists at least one initial state whose block-Kraft sum does not exceed unity. The stronger result is that \ell9 grows not even subexponentially but is uniformly bounded:

s2[1+log(1+α/s2)]s^2[1+\log(1+\alpha^\ell/s^2)]0

with row and total sums bounded by s2[1+log(1+α/s2)]s^2[1+\log(1+\alpha^\ell/s^2)]1 and s2[1+log(1+α/s2)]s^2[1+\log(1+\alpha^\ell/s^2)]2 respectively. The proof uses the Perron eigenvector ratio s2[1+log(1+α/s2)]s^2[1+\log(1+\alpha^\ell/s^2)]3 together with the existence of a path of length at most s2[1+log(1+α/s2)]s^2[1+\log(1+\alpha^\ell/s^2)]4 whose entries are each at least s2[1+log(1+α/s2)]s^2[1+\log(1+\alpha^\ell/s^2)]5. A further structural argument shows that reducible machines confer no asymptotic advantage even for individual sequences: along any infinite sequence, the set of infinitely visited states is closed and induces an irreducible machine on fewer states.

Converse bounds on compression and prediction

The uniform bound of Theorem 2 yields converse lower bounds that improve on what Ziv–Lempel's Lemma 2 provides. Via Jensen's inequality applied to the bounded Kraft sums,

s2[1+log(1+α/s2)]s^2[1+\log(1+\alpha^\ell/s^2)]6

and for stationary sources the bound can be sharpened using conditional entropies. The penalty term now vanishes at rate s2[1+log(1+α/s2)]s^2[1+\log(1+\alpha^\ell/s^2)]7, versus the s2[1+log(1+α/s2)]s^2[1+\log(1+\alpha^\ell/s^2)]8 rate implied by the Ziv–Lempel inequality or by the linear-in-s2[1+log(1+α/s2)]s^2[1+\log(1+\alpha^\ell/s^2)]9 bound available without irreducibility. For individual sequences, a cyclic-extension argument produces an analogous bound in terms of shift-invariant empirical conditional entropies, which via Ziv's inequality connects to LZ complexity s=1s=10, with the dominant residual term s=1s=11 after optimizing over s=1s=12 proportional to s=1s=13.

For prediction, the paper considers s=1s=14-state FS predictors with additive loss s=1s=15 on a group alphabet and constructs an auxiliary Gibbs-type model s=1s=16. Comparing an upper bound on the codelength of a Shannon code for this model against the GKI-based lower bound yields

s=1s=17

where s=1s=18. The bound is informative when s=1s=19 and =1\ell=10, and is tight essentially for sequences generated as prediction output plus i.i.d. noise distributed according to the Gibbs marginal.

Extension to side information

When SI =1\ell=11 is available to both encoder and decoder, each SI symbol induces its own Kraft matrix =1\ell=12, and feasibility is governed by the joint spectral radius (JSR) of the family =1\ell=13: IL implies =1\ell=14. A verifiable sufficient condition is the existence of a common positive sub-invariant vector =1\ell=15 with =1\ell=16 for all =1\ell=17; the all-one vector case corresponds to per-state Kraft sums bounded by one for every SI sequence. Crucially, individual spectral radii being at most one is not sufficient: the paper exhibits matrices =1\ell=18 and =1\ell=19 with E=(X,Y,Z,f,g)E=(\mathcal{X},\mathcal{Y},\mathcal{Z},f,g)0 arbitrarily small yet E=(X,Y,Z,f,g)E=(\mathcal{X},\mathcal{Y},\mathcal{Z},f,g)1 on the order of E=(X,Y,Z,f,g)E=(\mathcal{X},\mathcal{Y},\mathcal{Z},f,g)2. Since exact JSR computation is undecidable even for rational nonnegative matrices, the result should be read as a structural constraint, with the rich literature of JSR upper/lower bounds serving as practical tools.

Extension to lossy compression

For lossy coding, the model maps each source block E=(X,Y,Z,f,g)E=(\mathcal{X},\mathcal{Y},\mathcal{Z},f,g)3 to a reproduction E=(X,Y,Z,f,g)E=(\mathcal{X},\mathcal{Y},\mathcal{Z},f,g)4 within distortion budget, then compresses losslessly with an IL FS encoder. Because multiple sources map to the same reproduction, the effective Kraft matrix satisfies E=(X,Y,Z,f,g)E=(\mathcal{X},\mathcal{Y},\mathcal{Z},f,g)5 entry-wise, where E=(X,Y,Z,f,g)E=(\mathcal{X},\mathcal{Y},\mathcal{Z},f,g)6, giving

E=(X,Y,Z,f,g)E=(\mathcal{X},\mathcal{Y},\mathcal{Z},f,g)7

in the spirit of Campbell-type lossy Kraft inequalities. For additive distortion measures, E=(X,Y,Z,f,g)E=(\mathcal{X},\mathcal{Y},\mathcal{Z},f,g)8 can be estimated by the method of types, Chernoff bounds, or saddle-point integration.

Limitations and open questions

Several caveats bear directly on the strength of the results. First, whether E=(X,Y,Z,f,g)E=(\mathcal{X},\mathcal{Y},\mathcal{Z},f,g)9 is sufficient for information losslessness remains open in general; only restricted affirmative answers (per-state Kraft satisfaction, or the block-level prefix-code relaxation) are given. Second, the converse bounds carry explicit penalties involving s×ss\times s0 and the number of states, so their tightness depends on regimes where s×ss\times s1 dominates these quantities; the individual-sequence bound retains a s×ss\times s2 slack from Ziv's inequality. Third, the side-information extension is a structural condition only — undecidability of the JSR precludes a computational criterion, and the common sub-invariant vector condition is sufficient but not necessary. Finally, the lossy extension inherits the looseness of bounding s×ss\times s3 by s×ss\times s4.

Conclusion

The paper replaces Ziv–Lempel's asymptotic, super-alphabet-based generalized Kraft inequality with an exact single-letter condition: information losslessness of a finite-state encoder is equivalent (as a necessary condition) to spectral radius at most one of an explicitly defined Kraft matrix. In the irreducible case, Perron–Frobenius theory yields uniformly bounded matrix powers and s×ss\times s5 converse penalties, improving both probabilistic and individual-sequence lower bounds for compression and prediction. Extensions to side information and lossy coding recast feasibility in terms of joint spectral radii and type-class sizes respectively, while leaving sufficiency of the spectral-radius condition as the principal open question.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.