- The paper presents a new spectral-radius condition substituting Ziv–Lempel's 1978 asymptotic super-alphabet-based generalized Kraft inequality, showing that information losslessness of a finite-state encoder is equivalent to the spectral radius of a Kraft matrix.
- The key finding demonstrates that the spectral radius of the Kraft matrix is a necessary condition that is proven to be equivalent to the classical Kraft inequality when the number of states is one, thereby providing a unified framework for both lossless and lossy coding.
- The paper shows that reducible encoders confer no asymptotic advantage, the Craft matrix preserves topological probabilistic eigenvalues bounded on reducible and irreducible graphs, implying a uniformly bounded $2^{(s-1)L_{ ext{max}}$ for $K^n$.
Motivation and relation to prior work
Kraft's inequality, extended by McMillan to uniquely decodable codes, is the classical necessary and sufficient condition on codeword lengths for variable-length lossless codes, and underpins the converse to the source coding theorem. When the encoder has memory — modeled as a finite-state (FS) encoder in the sense of Ziv and Lempel (2601.16594), whose output and next state are determined recursively from the current symbol and state via functions f and g — the scalar Kraft inequality no longer applies directly. The only prior generalization for information-lossless (IL) FS encoders appears as Lemma 2 of Ziv–Lempel's 1978 paper, which bounds the block-level Kraft sum over ℓ-vectors by s2[1+log(1+αℓ/s2)]. The paper identifies two defects of this formulation: it does not recover the classical Kraft inequality when s=1 (the right-hand side exceeds unity even for ℓ=1), and it is posed only at the level of super-alphabet extensions rather than at the single-symbol/state level at which the encoder is defined.
The Kraft matrix and the spectral-radius condition
The central construction associates to any IL FS encoder E=(X,Y,Z,f,g) an s×s nonnegative Kraft matrix K, with entries
Kzz′=∑{x: g(z,x)=z′}2−L[f(z,x)].
The main theorem asserts that information losslessness implies g0, i.e., no eigenvalue of g1 may have modulus exceeding one. The proof shows that each entry of g2 equals g3, which the IL property (injectivity of g4 given fixed endpoints) bounds linearly in g5, namely by g6. If instead g7, Perron–Frobenius theory supplies a nonnegative right eigenvector g8, domination of g9 by a multiple of ℓ0, and hence exponential growth of ℓ1 — contradicting the linear bound. This condition reduces exactly to the classical Kraft inequality for ℓ2, since ℓ3 degenerates to the ordinary Kraft sum.
An illustrative three-state implementation of the length-2 block code ℓ4 has eigenvalues ℓ5: the row sum corresponding to the start state exceeds unity, demonstrating that a general IL FS encoder need not satisfy per-state prefix-code Kraft conditions — the FS model is strictly more general than a machine implementing a separate UD code at every state. Whether ℓ6 is also sufficient for information losslessness is left open; sufficiency holds trivially if every state satisfies its own Kraft inequality, and at the block level one can always prepend ℓ7 bits to obtain a genuine prefix code.
Irreducible encoders: bounded Kraft sums
For irreducible encoders — those whose adjacency graph is strongly connected — the paper derives several equivalent Collatz–Wielandt forms of the GKI, including the statement that for every ℓ8 there exists at least one initial state whose block-Kraft sum does not exceed unity. The stronger result is that ℓ9 grows not even subexponentially but is uniformly bounded:
s2[1+log(1+αℓ/s2)]0
with row and total sums bounded by s2[1+log(1+αℓ/s2)]1 and s2[1+log(1+αℓ/s2)]2 respectively. The proof uses the Perron eigenvector ratio s2[1+log(1+αℓ/s2)]3 together with the existence of a path of length at most s2[1+log(1+αℓ/s2)]4 whose entries are each at least s2[1+log(1+αℓ/s2)]5. A further structural argument shows that reducible machines confer no asymptotic advantage even for individual sequences: along any infinite sequence, the set of infinitely visited states is closed and induces an irreducible machine on fewer states.
Converse bounds on compression and prediction
The uniform bound of Theorem 2 yields converse lower bounds that improve on what Ziv–Lempel's Lemma 2 provides. Via Jensen's inequality applied to the bounded Kraft sums,
s2[1+log(1+αℓ/s2)]6
and for stationary sources the bound can be sharpened using conditional entropies. The penalty term now vanishes at rate s2[1+log(1+αℓ/s2)]7, versus the s2[1+log(1+αℓ/s2)]8 rate implied by the Ziv–Lempel inequality or by the linear-in-s2[1+log(1+αℓ/s2)]9 bound available without irreducibility. For individual sequences, a cyclic-extension argument produces an analogous bound in terms of shift-invariant empirical conditional entropies, which via Ziv's inequality connects to LZ complexity s=10, with the dominant residual term s=11 after optimizing over s=12 proportional to s=13.
For prediction, the paper considers s=14-state FS predictors with additive loss s=15 on a group alphabet and constructs an auxiliary Gibbs-type model s=16. Comparing an upper bound on the codelength of a Shannon code for this model against the GKI-based lower bound yields
s=17
where s=18. The bound is informative when s=19 and ℓ=10, and is tight essentially for sequences generated as prediction output plus i.i.d. noise distributed according to the Gibbs marginal.
When SI ℓ=11 is available to both encoder and decoder, each SI symbol induces its own Kraft matrix ℓ=12, and feasibility is governed by the joint spectral radius (JSR) of the family ℓ=13: IL implies ℓ=14. A verifiable sufficient condition is the existence of a common positive sub-invariant vector ℓ=15 with ℓ=16 for all ℓ=17; the all-one vector case corresponds to per-state Kraft sums bounded by one for every SI sequence. Crucially, individual spectral radii being at most one is not sufficient: the paper exhibits matrices ℓ=18 and ℓ=19 with E=(X,Y,Z,f,g)0 arbitrarily small yet E=(X,Y,Z,f,g)1 on the order of E=(X,Y,Z,f,g)2. Since exact JSR computation is undecidable even for rational nonnegative matrices, the result should be read as a structural constraint, with the rich literature of JSR upper/lower bounds serving as practical tools.
Extension to lossy compression
For lossy coding, the model maps each source block E=(X,Y,Z,f,g)3 to a reproduction E=(X,Y,Z,f,g)4 within distortion budget, then compresses losslessly with an IL FS encoder. Because multiple sources map to the same reproduction, the effective Kraft matrix satisfies E=(X,Y,Z,f,g)5 entry-wise, where E=(X,Y,Z,f,g)6, giving
E=(X,Y,Z,f,g)7
in the spirit of Campbell-type lossy Kraft inequalities. For additive distortion measures, E=(X,Y,Z,f,g)8 can be estimated by the method of types, Chernoff bounds, or saddle-point integration.
Limitations and open questions
Several caveats bear directly on the strength of the results. First, whether E=(X,Y,Z,f,g)9 is sufficient for information losslessness remains open in general; only restricted affirmative answers (per-state Kraft satisfaction, or the block-level prefix-code relaxation) are given. Second, the converse bounds carry explicit penalties involving s×s0 and the number of states, so their tightness depends on regimes where s×s1 dominates these quantities; the individual-sequence bound retains a s×s2 slack from Ziv's inequality. Third, the side-information extension is a structural condition only — undecidability of the JSR precludes a computational criterion, and the common sub-invariant vector condition is sufficient but not necessary. Finally, the lossy extension inherits the looseness of bounding s×s3 by s×s4.
Conclusion
The paper replaces Ziv–Lempel's asymptotic, super-alphabet-based generalized Kraft inequality with an exact single-letter condition: information losslessness of a finite-state encoder is equivalent (as a necessary condition) to spectral radius at most one of an explicitly defined Kraft matrix. In the irreducible case, Perron–Frobenius theory yields uniformly bounded matrix powers and s×s5 converse penalties, improving both probabilistic and individual-sequence lower bounds for compression and prediction. Extensions to side information and lossy coding recast feasibility in terms of joint spectral radii and type-class sizes respectively, while leaving sufficiency of the spectral-radius condition as the principal open question.