---
title: 'Endpoint Decodability: Theory and Applications'
url: https://www.emergentmind.com/topics/endpoint-decodability
type: topic
---

# Endpoint Decodability: Theory and Applications

Searching arXiv for recent and foundational papers using “endpoint decodability” across coding, representation learning, and generative modeling.
Endpoint decodability is a domain-dependent notion of recoverability at an endpoint of a decoding, inference, or communication process. In the cited literature, it appears in at least five technically distinct forms: as sharp list-decoding behavior at the endpoint of a rate or radius trade-off in classical and quantum coding; as local decodability constraints for variable-length compression of sparse sources; as the requirement that intermediate neural activations decode back to the original input; as algebraic recoverability of the clean sample \(x_0\) from an intermediate diffusion or flow-matching state; and as linear recoverability of desired substreams at each receiving endpoint in MIMO coded caching [1010.3312, 1611.06673, 1504.02063, 2106.00769, 2607.06114, 2603.06534].

## 1. Scope and recurrent structure

A common misconception is that endpoint decodability denotes a single canonical property. The available literature instead uses the expression in several non-equivalent ways. In coding theory, “endpoint” often refers to a sharp boundary of a list-decoding trade-off, such as exact small-radius thresholds or rates approaching the Gilbert–Varshamov bound. In source coding, representation learning, diffusion, and MIMO delivery, the term is operational: a distinguished endpoint object must be recoverable from restricted probes, hidden states, noisy trajectory points, or linearly mixed streams.

| Setting | Endpoint notion | Formal criterion |
|---|---|---|
| List decoding | Sharp trade-off endpoint | \(A'(n,d,e)\), or \(R=1-H_q(\delta)-\epsilon\) |
| Sparse source coding | Local recovery of \(x_j\) | At most \(d\) probes to the codeword |
| Decodable neural networks | Input recovery from \(h_l\) | \(g_\phi(h_l)\) reconstructs in input space |
| Diffusion / flow matching | Clean-sample recovery from \((x_t,u_t)\) | \(\Delta_t\neq 0\) and \(x_0=\frac{\sigma_tu_t-\dot\sigma_t x_t}{\Delta_t}\) |
| MIMO coded caching | Linear recovery of desired streams | Conditions \(C_1\) and \(C_2\) |

Taken together, these uses suggest a broad common pattern: an intermediate representation is useful only insofar as it preserves enough structure to recover a target endpoint with explicit guarantees. The technical mechanisms, however, vary sharply across the cited works [1010.3312, 1611.06673, 1504.02063, 2106.00769, 2607.06114, 2603.06534].

## 2. Coding-theoretic endpoints: list size, radius, and rate

In coding theory, endpoint decodability is tied to extremal list-decoding thresholds. The paper of Jin–Xing–Zhang proves that random Euclidean self-orthogonal codes achieve the classical Gilbert–Varshamov bound “to the endpoint.” For any fixed prime power \(q\), any \(\delta\in(0,1-1/q)\) satisfying \(1-H_q(\delta)\le \tfrac12\), and any \(\epsilon>0\), a uniformly random Euclidean self-orthogonal code \(C\subseteq \mathbb F_q^n\) of rate
\[
R=\frac{\dim C}{n}=1-H_q(\delta)-\epsilon
\]
is, with probability \(1-q^{-n}\), \((\delta,L)\)-list-decodable for
\[
L=O\!\bigl(\tfrac1\epsilon\bigr)
\]
as \(n\to\infty\). The same framework extends to symplectic dual-containing codes, yielding the quantum rate
\[
R = 1-H_q(\delta)-\delta\log_q(q+1)-\epsilon,
\]
again with \((\delta,O(1/\epsilon))\)-list-decodability and probability \(1-q^{-n}\) [1611.06673].

The analysis uses the \(q\)-ary entropy function
\[
H_q(\delta)=\delta\log_q(q-1)-\delta\log_q\delta-(1-\delta)\log_q(1-\delta)
\]
and the asymptotic Hamming-ball volume
\[
\bigl|\{\,y\in\mathbb F_q^n : d_H(x,y)\le \delta n\,\}\bigr|
=\exp_q\bigl(nH_q(\delta)+o(n)\bigr).
\]
Its proof structure combines a union bound over bad lists, the Guruswami–Håstad–Kopparty limited-correlation lemma, and a counting lemma for self-orthogonal subspaces. For linearly independent \(v_1,\dots,v_t\in\mathbb F_q^n\) with \(t<k<n/2\), the fraction of \([n,k]\) self-orthogonal codes containing \(\{v_1,\dots,v_t\}\) is at most
\[
q^{\,\bigl((k+t-n-2)\,t+4k-1\bigr)}.
\]
This counting control, together with the probabilistic bound
\[
\Pr\bigl\{\,|\mathrm{span}(X_1,\dots,X_t)\cap B_n(0,\delta)|>Mt\,\bigr\}<q^{-\Omega(n)},
\]
drives the final failure estimate
\[
\Pr(\text{failure})<q^{-n}.
\]

A different endpoint notion appears in the exact small-radius theory of binary list decoding. There,
\[
A'(n,d,e)
\]
is defined as the smallest \(\ell\) such that every binary \((n,d)\)-code is \((e,\ell)\)-list-decodable; equivalently,
\[
A'(n,d,e)=\max\{|S|: S\subseteq \mathbb F_2^n \text{ is an } (n,d)\text{-code and }S\subseteq B(x,e)\text{ for some }x\}.
\]
For all \(d\ge 2e-3\), exact closed-form expressions are obtained. The principal regimes are
\[
A'(n,d,e)=1 \quad \text{if } d\ge 2e+1,
\]
\[
A'(n,2e,e)=\lfloor n/e\rfloor,
\]
\[
A'(n,2e-1,e)=\lfloor (n+1)/e\rfloor,
\]
\[
A'(n,2e-2,e)
=\max_{\alpha\in\{0,1\},\,\beta\ge 0,\,\alpha\beta=0}
\left\{\alpha+\beta+A\!\bigl(n-\alpha(e-2)-\beta(e-1),\,2e-2,\,e\bigr)\right\},
\]
and
\[
A'(n,2e-3,e)=A'(n+1,2e-2,e).
\]
For \(e\le 4\), exact values are determined, apart from \(42\) exceptional values of \(n\) when \(e=4\) [1010.3312].

These two coding-theoretic lines use the word “endpoint” differently. In one case it denotes asymptotic achievability of the rate boundary \(R=1-H_q(\delta)\); in the other, it denotes exact threshold points on the radius/list-size curve \(\{(e,\ell(e))\}\) with \(\ell(e)=A'(n,d,e)\). This suggests that, within coding theory, endpoint decodability is primarily a boundary-achievability concept rather than an inversion problem.

## 3. Variable-length sparse compression under local decodability constraints

Pananjady and Courtade study a variable-length source-coding problem for \(r\)-sparse binary sequences under local decodability constraints. The source \(X^n\) is chosen uniformly from all \(r\)-sparse binary vectors in \(\{0,1\}^n\), with
\[
\mathrm{supp}(X^n)\triangleq\{i:X_i=1\}, \qquad |\mathrm{supp}(X^n)|=r.
\]
A variable-length encoder
\[
c:\binom{[n]}{r}\to\{0,1\}^*
\]
must be invertible, and a code is an \((r,d,n)\) code if there exists a decoder which, on input \((j,\ell)\), can recover \(x_j\) by probing at most \(d\) bits of the codeword. The model distinguishes adaptive access, where each probe location may depend on previous probe outcomes, from non-adaptive access, where all \(d\) probe locations are chosen in advance as a function only of \(j\) and \(\ell\) [1504.02063].

The main quantity is the average blocklength
\[
L=\mathbb E[\ell(c(X^n))].
\]
For adaptive schemes, the lower bound is
\[
\mathbb E[\ell(c(X^n))]+1 \ge \frac{r\,d+1}{4e}\cdot\left(\binom{n}{r}^{1/(r\,d+1)}-1\right).
\]
As \(d\to\infty\) with \(n,r\) fixed, the right-hand side tends to \(\frac{1}{4e}\ln \binom nr\), recovering the usual entropy bound up to a constant. For non-adaptive schemes, there exists an \((r,d,n)\) code with
\[
\mathbb E[\ell(c(X^n))]
\le 30\cdot(r\,d+1)\cdot\bigl((r+1)^{(r+1)}\binom nr\bigr)^{1/(r\,d+1)}.
\]
When
\[
\log \binom nr=\Omega(r\,d)
\qquad\text{and}\qquad
d=\Omega(\log r),
\]
these bounds coincide in order:
\[
\mathbb E[\ell(c(X^n)])=\Theta\bigl(r\,d\cdot \binom nr^{1/(r\,d+1)}\bigr).
\]

The converse generalizes the Sperner/LYM-lemma method from static bit-probe lower bounds. For each \(x^n\) and each codeword length \(k\), one records the set of at most \(r\,d\) “probed positions + bit values” needed to recover the \(r\) ones of \(x^n\). These sets form an antichain in \(\{1,\dots,k\}\times\{0,1\}\), and the LYM inequality bounds the number of short codewords. The achievability proof uses a random-codebook construction: for each \(k\ge r\,d+1\), choose a random subset \(S_k\subset [n]\) of size approximately
\[
\frac{r}{r+1}\cdot \frac{\binom{k}{d}}{\binom{r\,d}{d}},
\]
assign random \(d\)-subsets \(T_{j,k}\subset [k]\), encode by the indicator of \(\bigcup_{i\in\mathrm{supp}(x^n)}T_{i,k}\), and decode by probing \(T_{j,k}\) and returning the AND.

Because the lower bound already applies to adaptive decoders while the upper bound is non-adaptive, the asymptotic exponent in \(n\) is the same in both models in most regimes. In particular, for fixed \(r,d\),
\[
L=\Theta(n^{r/(r\,d+1)}),
\]
while for \(r=o(n)\) and \(d=\Theta(\log n)\),
\[
L=\Theta(r\,d)=\Theta\!\left(\log \binom nr\right).
\]
The same construction also yields a one-round SpeedLimit communication protocol for the membership function \(f(i,S)=1_{i\in S}\), with average speed-limit
\[
\mathbb E[2^z]
\le 30(r\,d+1)\bigl((r+1)^{(r+1)}\binom nr\bigr)^{1/(r\,d+1)}.
\]

In this literature, endpoint decodability means that arbitrary source coordinates remain queryable after compression, even when the codeword has variable length and only \(d\) probes are permitted.

## 4. Decodable neural networks and activation-to-input inversion

In the neural-network setting, endpoint decodability is defined directly on intermediate activations. Let \(f_\theta\) be a neural-network encoder with intermediate activations \(h_1,\dots,h_L\), where
\[
h_l=f^{(l)}_\theta(x).
\]
An activation \(h_l\) is endpoint-decodable if there exists a decoder \(g_\phi\) such that \(g_\phi(h_l)\) reconstructs back into the original input space. Writing
\[
S(h)=\{x': f_\theta(x')=h\}
\]
for the preimage set, the target behavior is
\[
g_\phi(h_l)\sim p(x\mid f_\theta(x)=h_l),
\]
implemented in practice by maximizing \(\log p(x\mid g_\phi(h_l))\) over training data [2106.00769].

The joint training objective for a Decodable Neural Network (DecNN) is
\[
L_{\mathrm{total}}(\theta,\phi;x,y)
=
L_{\mathrm{cls}}(f_\theta(x),y)
+
\beta\cdot \frac{1}{L}\sum_{l=1}^L
L_{\mathrm{rec}}(g_\phi(h_l),x),
\]
where \(L_{\mathrm{cls}}\) is the cross-entropy \(-\log p(y\mid f_\theta(x))\), \(L_{\mathrm{rec}}\) is for example \(\log p(x\mid g_\phi(h_l))\) or MSE for a Gaussian decoder, and \(\beta>0\) weights the decoding regularizer. The decoder \(g_\phi\) is a ResNet-18–style decoder with upsampling layers, shared across all layers. Backpropagation through both \(g_\phi\) and \(f_\theta\) forces representations to remain simultaneously discriminative and invertible.

Because each \(g_\phi(h_l)\) has the same shape as the original input, the model can be recursively self-composed. Repeating the map \(f_\theta\circ g_\phi\) to depth \(D\) yields a tree of implicit classifiers of size \(O(L^D)\). The recursive objective is
\[
\hat L(\theta,\phi;x,y)
= \log p(y\mid f_\theta(x))
+ \sum_{d=1}^D \alpha^d\Bigl[
\log p\bigl(y\mid f_\theta(g_\phi(h^{(d-1)}_{l_{d-1}}))\bigr)
+ \beta \log p\bigl(x\mid g_\phi(h^{(d-1)}_{l_{d-1}})\bigr)
\Bigr],
\]
where \(l_{d-1}\in\{1,\dots,L\}\) is sampled at each depth and \(\alpha\in(0,1]\) down-weights deeper compositions. At test time, \(N\) random paths produce an ensemble prediction, and the uncertainty is the ensemble entropy
\[
\mathrm{Uncertainty}(x)
=
-\sum_{c=1}^K \hat p(c)\log \hat p(c),
\]
with \(\hat p(c)\) the fraction of sampled paths predicting class \(c\).

Empirically, decoding does not degrade classification accuracy. On MNIST, FashionMNIST, and CelebA, a standard MLP attains \(97.1\), \(87.7\), and \(90.8\); DecNN attains \(97.8\), \(88.8\), and \(90.8\); and ReDecNN attains \(97.5\), \(88.1\), and \(90.8\). For misclassification detection, ReDecNN yields ROC-AUC \(0.869\), \(0.795\), and \(0.692\), compared with MC-Dropout \(0.878\), \(0.818\), and \(0.657\), and naive ensemble \(0.531\), \(0.559\), and \(0.605\). For calibration, Expected Calibration Error drops from \(1.18\), \(0.56\), and \(2.75\) for standard networks to \(0.33\), \(0.21\), and \(1.92\) for ReDecNN. The method also incurs approximately \(6\times\) slower training than standard MLPs, and its success depends on the quality of the generative decoder \(g_\phi\) [2106.00769].

In this setting, endpoint decodability is neither a combinatorial threshold nor a communication constraint. It is an architectural regularity condition: each hidden layer must remain rich enough to decode back into the input space.

## 5. Diffusion and flow matching: decoding the clean endpoint from intermediate states

In diffusion and flow-matching models, endpoint decodability is an algebraic property of affine probability paths
\[
x_t=\alpha_t x_0+\sigma_t \epsilon,\qquad \epsilon\sim \mathcal N(0,I),\qquad t\in[0,1],
\]
with smooth schedules \(\alpha_t,\sigma_t\) satisfying \(\alpha_0=0,\sigma_0=1\) and \(\alpha_1=1,\sigma_1=0\). Differentiation gives the conditional velocity
\[
u_t=\frac{dx_t}{dt}=\dot \alpha_t x_0+\dot \sigma_t \epsilon,
\]
so \((x_t,u_t)\) obey the \(2\times 2\) linear system
\[
\begin{bmatrix}x_t \\ u_t\end{bmatrix}
=
\begin{bmatrix}
\alpha_t & \sigma_t\\
\dot\alpha_t & \dot\sigma_t
\end{bmatrix}
\begin{bmatrix}
x_0\\ \epsilon
\end{bmatrix}.
\]
The path is endpoint-decodable at time \(t\) if
\[
\Delta_t
=
\det
\begin{pmatrix}
\alpha_t & \sigma_t\\
\dot\alpha_t & \dot\sigma_t
\end{pmatrix}
=
\dot\alpha_t\sigma_t-\alpha_t\dot\sigma_t
\neq 0.
\]
The paper states that almost all standard schedules, including VP/VE diffusion, EDM, and linear FM, satisfy \(\Delta_t\neq 0\) for \(t>0\) [2607.06114].

When \(\Delta_t\neq 0\), Cramer’s rule yields the endpoint decoder
\[
x_0=\frac{\sigma_tu_t-\dot\sigma_t x_t}{\Delta_t}.
\]
If \(u_t\) is approximated by a learned velocity \(v_\theta(x_t,t)\), the induced predictor is
\[
\hat x_0(x_t,t)=\frac{\sigma_t v_\theta(x_t,t)-\dot\sigma_t x_t}{\Delta_t}.
\]
Under the standard mean-squared-error objective
\[
L=\mathbb E_{t,x_0,\epsilon}\bigl[\|v_\theta(x_t,t)-u_t\|^2\bigr],
\]
the Bayes-optimal velocity is \(v^*(x,t)=\mathbb E[u_t\mid x_t=x]\), and the same algebraic decoder recovers the posterior mean:
\[
\frac{\sigma_t v^*(x,t)-\dot\sigma_t x}{\Delta_t}
=
\mathbb E[x_0\mid x_t=x].
\]
Thus the endpoint decoder coincides with the MMSE estimator of the clean sample, even when the model is not explicitly trained to predict \(x_0\).

The error decomposition is
\[
\mathbb E[\|\hat x_0-x_0\|^2]
=
\mathbb E[\|e_t\|^2]+U(t),
\]
where \(U(t)=\mathbb E_{x_t}[\mathrm{Tr}\,\mathrm{Var}(x_0\mid x_t)]\) is the irreducible uncertainty and \(\hat x_0=m_t(x_t)+e_t(x_t)\) with \(m_t=\mathbb E[x_0\mid x_t]\). Crucially, no term involves \(\ddot\alpha_t\) or \(\ddot\sigma_t\). The paper therefore characterizes endpoint prediction as curvature-independent and argues that trajectory straightness is sufficient but not necessary for acceleration.

This observation leads to Truncated Jump Sampling (TJS): stop the ODE at an early-exit time \(t^*\) and return the decoded \(x_0\). With total steps \(K\) and early-exit fraction \(\gamma\in(0,1]\), set \(k^*=\lceil \gamma K\rceil\), run the sampler only to \(k^*\), and return \(\hat x_0(x,t^*)\). Total NFEs equal \(k^*+1\), since the final decode costs one network call. The quality-speed trade-off is empirically concave, and a practical rule of thumb is \(\gamma\approx 0.5\)–\(0.7\), recovering at least \(95\%\) of full-ODE quality while saving \(30\%\)–\(50\%\) NFEs.

The reported experiments span SDXL, SD3.5M, Z-Image-Turbo, and three class-conditional benchmarks. On ImageNet-256 with CFG \(=1.0\), full \(40\)-step FID is \(17.52\); TJS-\(0.5\) with \(21\) NFE gives FID \(44.91\); TJS-\(0.7\) with \(29\) NFE gives FID \(17.16\); and TJS-\(0.8\) with \(33\) NFE gives FID \(15.16\). For SDXL and SD3.5M on DrawBench, semantic metrics saturate by \(k^*\approx 12\), corresponding to \(57\%\) saving, while ImageReward converges by \(k^*\approx 18\), corresponding to \(37\%\) saving. On Z-Image-Turbo, \(k^*=2\) gives \(3\) NFE and \(70\%\) saving with all metrics at least \(95\%\). Across all six families, TJS yields \(20\%\)–\(70\%\) NFE reduction with near-matched quality, without retraining, distillation, or architecture change [2607.06114].

## 6. Linear decodability in MIMO coded caching

In the MIMO coded-caching setting, endpoint decodability refers to linear recoverability at each receiving endpoint. A single base station with \(L\) transmit antennas serves \(K\) users, each with \(G\) receive antennas. The coded-caching gain is
\[
t=\frac{KM}{N}.
\]
In time slot \(i\), the base station selects a target user subset \(U(i)\) of size \(\Omega\) and transmits
\[
x(i)=\sum_{T\subseteq U(i),\,|T|=t+1} v_T(i)\,X_T,
\]
where \(X_T\) is the XOR codeword intended for multicast group \(T\) and \(v_T(i)\in \mathbb C^{L\times \theta_T(i)}\) is the beamforming matrix. User \(k\in U(i)\) observes
\[
y_k(i)=H_k(i)\,x(i)+n_k(i),
\]
with \(H_k\in\mathbb C^{G\times L}\) and \(n_k\sim \mathcal{CN}(0,I)\) [2603.06534].

Slot \(i\) is linearly decodable at user \(k\) if there exists a receive-beamforming matrix \(U_k(i)\in \mathbb C^{G\times \beta_k(i)}\) such that
\[
U_k(i)^H H_k(i)\,v_T(i)=0
\]
for every stream not intended for \(k\), while the aggregated desired-stream matrix has full column rank \(\beta_k(i)\). No successive-interference cancellation is needed: each user zero-forces all unwanted streams in one shot and recovers its \(\beta_k(i)\) desired streams through a full-rank linear filter.

The central criterion is necessary and sufficient. For every slot \(i\) and every multicast index \(T\subseteq U(i)\), linear decodability at every \(k\in U(i)\) is possible if and only if
\[
C_1\qquad
\sum_{k'\in U(i)\setminus T}\beta_{k'}(i)+\theta_T(i)\le L,
\qquad \forall\,T\subseteq U(i),\, |T|=t+1,
\]
and
\[
C_2\qquad
\beta_k(i)\le G,
\qquad \forall\,k\in U(i).
\]
Condition \(C_2\) is the receive-antenna constraint. Condition \(C_1\) follows from rank-nullity after stacking the interference channels into
\[
\bar H_T(i)=\bigl[H_{k'}(i)\,U_{k'}(i)\bigr]_{k'\in U(i)\setminus T},
\]
whose nullspace must support the \(\theta_T(i)\) replicated substreams sharing the same multicast index.

The paper gives both symmetric and asymmetric examples. In the symmetric case with \(L=10\), \(G=3\), \(t=1\), \(\Omega=5\), \(\theta_T(i)=1\), and \(\beta_k(i)=2\) for all \(k\), any \(T\) of size \(2\) has \(3\) interfering users, so
\[
\sum_{k'\notin T}\beta_{k'}=3\times 2=6,
\qquad
6+1=7\le 10,
\]
and \(\beta_k=2\le 3\), hence \(C_1\) and \(C_2\) both hold. In the asymmetric case with \(L=11\), \(G=6\), \(t=2\), \(\Omega=5\), a symmetric reference \(\beta_k=3\), \(\theta_T=1\) yields \(\mathrm{DoF}_{\mathrm{ref}}=15\), while the decomposition-reassignment construction adds \(m=3\) extra multicast indices per slot, producing \(\sum_k\beta_k=24\) total DoF per slot. The worst-case check gives
\[
\sum_{k'\notin T}\beta_{k'}\le 9,
\qquad
9+2=11\le L,
\]
and \(\max \beta_k=4\le G=6\), so linear decodability is preserved [2603.06534].

The scheduling framework replicates the \(B\times S\) scheduling table \(\Delta\) times, decomposes and reassigns some of the last \(\Delta S-S\) columns, and chooses integers \((m,\Delta,S')\) such that
\[
m\cdot S'=B(\Delta S-S).
\]
The DoF increase is
\[
\Delta\mathrm{DoF}=m(t+1).
\]
A balanced greedy hypergraph-based algorithm controls pairwise overlap among new indices, enforces \(\beta_k(i)\le G\), and respects
\[
\theta_T(i)+\sum_{k'\notin T}\beta_{k'}\le L.
\]
The reported simulations show that asymmetric scheduling fills in gaps left by symmetric-only stream allocations and enlarges the feasible DoF region without violating linear decodability.

Taken together, the cited literature suggests that endpoint decodability is best understood as a family of endpoint-recovery or boundary-achievability principles rather than a single invariant. In coding theory it governs sharp extremal limits; in source coding it constrains probe complexity under compression; in neural architectures it enforces invertible hidden states; in diffusion and flow matching it becomes an algebraic decoder for \(x_0\); and in MIMO coded caching it reduces to rank-nullity conditions for one-shot linear recovery at each user endpoint.

Source: https://www.emergentmind.com/topics/endpoint-decodability