---
title: Function-Compression Capacity
url: https://www.emergentmind.com/topics/function-compression-capacity
type: topic
---

# Function-Compression Capacity

Function-compression capacity is a context-dependent technical notion for the efficiency with which a function, a function-defined task, or a function-generated family can be represented, communicated, or internally simplified under explicit correctness, distortion, or performance constraints. In current literature, the term does not denote a single invariant quantity: zero-error distributed source coding defines it as an asymptotic transmission cost for computing a target function [2508.02996]; network function computation often uses the reciprocal maximum computing rate [1710.02252]; graph-entropy formulations characterize the minimum rates needed to convey function-relevant colorings rather than raw sources [1011.5496]; task-oriented compression measures degradation of a downstream optimization or valuation objective [2405.07808, 2509.06810]; and operator-compression work treats singular-value decay or effective rank as an operational measure of how many basis elements are needed to represent a function family [2605.24814]. A plausible implication is that the expression is best treated as an umbrella term whose precise meaning is fixed by the compressed object, the admissible compression class, and the preservation criterion.

## 1. Scope, conventions, and competing formalizations

The most explicit uses of the phrase occur in zero-error communication settings, but even there the sign convention is not uniform. In distributed source coding for vector-linear functions, an admissible code \(\mathbf C\) has overall rate
\[
R(\mathbf C)=\frac{n(\mathbf C)}{k},
\]
and the function-compression capacity is
\[
\mathcal{C}(s,m,\Omega,f)=\inf\{R:\;R\text{ is achievable}\},
\]
so smaller values are better because the quantity is a normalized encoder-to-decoder transmission cost per function computation [2508.02996]. By contrast, network function computation defines computing capacity as the maximum asymptotic rate \(k/n\), so larger values are better [1710.02252]. The vector-linear source-coding model makes the reciprocity explicit through
\[
\mathcal{C}(s,m,\Omega,T)=\frac{1}{\mathcal{C}(\mathcal N,T)},
\]
linking minimum communication cost and maximum computing throughput [2508.02996].

A second exact convention appears in zero-error distributed compression of the binary arithmetic sum. There the compression capacity is the maximum average number of times the function can be compressed with zero error per one use of the system, again a throughput notion rather than a cost notion [2305.06619]. A third family of definitions replaces asymptotic communication rate by graph entropy, effective rank, or task loss. In graph-based functional compression, the relevant quantity is the minimum rate needed to transmit graph colorings that preserve the target function [1011.5496]. In imaginary-time Green-function compression, the decisive quantity is the effective rank
\[
N_{\rm eff}(\Lambda;\epsilon)=\#\{\,l\mid s_l/s_0>\epsilon\,\},
\]
which counts singular modes above a relative threshold [2605.24814]. In goal-oriented compression for power scheduling, the core quantity is the expected squared degradation in optimal utility,
\[
\Gamma=\mathbb E_\ell\!\left[\left|u(x^\star(\ell);\ell)-u(x^\star(\widehat\ell);\ell)\right|^2\right],
\]
not source reconstruction error [2405.07808].

The literature also contains nearby but non-equivalent uses of “function capacity.” “Estimating the Capacities of Function-as-a-Service Functions” defines Function Capacity as “the maximal number of concurrent invocations the function can serve in a time without violating the SLO,” which is a workload-capacity notion for serverless systems rather than a compression notion [2201.11454]. This terminological divergence is substantive rather than stylistic.

## 2. Zero-error distributed source coding and network function computation

In the distributed source-coding formulation for vector-linear functions, there are \(s\) sources, \(m\) encoders, one decoder, arbitrary source-to-encoder connectivity \(\Omega\), and a target function represented by an \(s\times r\) column-full-rank matrix \(T\) over \(\mathbb F_q\) [2508.02996]. For an admissible \(k\)-shot code, encoder \(j\) induces a required number of link uses
\[
n_j(\mathbf C)=\left\lceil \log_{|\mathcal A|} |\operatorname{Im}\varphi_j| \right\rceil,
\]
with overall rate determined by the worst encoder. The paper gives a general lower bound, valid for arbitrary connectivity states and arbitrary vector-linear functions with \(1<\operatorname{Rank}(T)<s\),
\[
\mathcal{C}(s,m,\Omega,T)\ge
\max_{C\in\Lambda(\mathcal N)}
\max_{\mathcal P_C}
\frac{\operatorname{rank}_{\mathcal P_C}(T)}{|C|},
\]
where \(\mathcal P_C\) ranges over strong partitions of cut sets. For the smallest nontrivial regime \(s=3\), \(\operatorname{Rank}(T)=2\), and \(m\le 3\), all \(3\times2\) column-full-rank matrices reduce to two capacity types,
\[
T_1=\begin{bmatrix}1&0\\0&1\\1&1\end{bmatrix},
\qquad
T_2=\begin{bmatrix}1&0\\1&0\\0&1\end{bmatrix},
\]
and the capacities are fully characterized; the only exceptional \(T_2\) cases have exact capacity \(3/4\), showing that the general lower bound is not always tight [2508.02996].

The binary arithmetic-sum problem yields a complete closed-form classification for all four encoder-observation patterns [2305.06619]. With \(f(x,y)=x+y\), \(C_1\ge C_2\), and switches \(s_1,s_2\) controlling encoder side information, the exact capacities are
\[
C(00;C_1,C_2;f)=C_2,\qquad
C(10;C_1,C_2;f)=C_2,
\]
\[
C(11;C_1,C_2;f)=\frac{C_1+C_2}{\log 3},
\qquad
C(01;C_1,C_2;f)=(C_1-C_2)\log_3 2 + C_2.
\]
The asymmetric \((01;C_1,C_2;f)\) case is the nontrivial one. Its converse is based on a conflict graph \(G(M,L)\) whose chromatic number equals the sumset size \(|M+L|\), together with the lower bound
\[
|\mathcal A^k+L|\ge 2^k |L|^{\log 3-1}.
\]
That exponent arises from the paper’s “aitch-function” \(h(\ell)=\ell^{\log 3-1}\) [2305.06619].

Network function computation uses the reciprocal viewpoint. For a directed acyclic network \(\mathcal N=(G,S,\rho)\) and target function \(f\), the computing capacity is the maximum asymptotic rate at which the sink can compute the function with zero error [1710.02252]. The paper introduces an improved upper bound
\[
\mathcal C(\mathcal N,f)\le
\min_{C\in\Lambda(\mathcal N)}
\frac{|C|}{\log_{|\mathcal A|} n_{C,f}},
\]
where \(n_{C,f}\) is defined through strong cut partitions and counts of locally realizable tuples of partition-induced equivalence classes. This bound is tight for the arithmetic-sum non-tree example, yielding
\[
\mathcal C(\mathcal N,f)=\frac{2}{1+\log_2 3},
\]
but it is not generally achievable: for the reverse butterfly network with binary maximum, the improved upper bound equals \(2\), yet the paper proves strict non-achievability [1710.02252]. A plausible implication is that cutwise distinguishability counts do not exhaust the global combinatorics of zero-error function computation.

## 3. Graph entropy, valid colorings, and network functional compression

A second major tradition treats function-compression capacity as a graph-entropy rate problem. For a function \(f(X_1,\dots,X_k)\), the characteristic graph \(G_{X_i}\) connects two values of \(X_i\) when they must be distinguished because there exists a compatible assignment of the other variables that changes the function value [1011.5496]. The single-source rate object is graph entropy,
\[
H_{G_{X_1}}(X_1)=\min_{X_1\in W_1\in\Gamma(G_{X_1})} I(X_1;W_1),
\]
and the conditional version is
\[
H_{G_{X_1}}(X_1|X_2)=
\min_{\substack{X_1\in W_1\in\Gamma(G_{X_1})\\ W_1-X_1-X_2}}
I(W_1;X_1|X_2),
\]
which reduces to \(H(X_1|X_2)\) when the target function is the identity [1011.5496]. Power graphs and \(\epsilon\)-colorings provide the asymptotic bridge from combinatorial colorings to achievable compression rates.

For the depth-one tree with correlated sources, the paper gives the exact rate region by replacing Slepian–Wolf entropies with graph-entropic terms [1011.5496]. In the two-source case,
\[
R_{11}\ge H_{G_{X_1}}(X_1|X_2),\qquad
R_{12}\ge H_{G_{X_2}}(X_2|X_1),
\]
\[
R_{11}+R_{12}\ge H_{G_{X_1},G_{X_2}}(X_1,X_2).
\]
For general tree networks, it derives stagewise lower bounds in terms of conditional graph entropies of the subtree variables, and for independent sources the same lower bound is achieved arbitrarily closely by a recursive scheme in which intermediate nodes decode incoming colors, identify an independent set in their own characteristic graph, and transmit a minimum-entropy coloring of that induced functional state [1011.5496]. Because graph entropy does not satisfy the chain rule, relaying is not generally optimal; the paper isolates “chain-rule proper sets” as the special case in which relaying suffices.

The central structural constraint is the Coloring Connectivity Condition (C.C.C.). Individually valid source colorings are insufficient if a joint color tuple can correspond to multiple disconnected support components with different function values. C.C.C. requires each joint coloring class either to be connected or, if disconnected, to have the same function value on all components [1011.5496]. The paper proves that C.C.C. is both necessary and sufficient for coloring-based decodability, and that any achievable coding scheme induces \(\epsilon\)-colorings satisfying C.C.C. This makes C.C.C. the exact graph-theoretic criterion that separates merely valid local colorings from globally function-preserving colorings.

Helper-based work generalizes the graph viewpoint by introducing functional common information. “Applications of Common Information to Computing Functions” defines a nested helper variable \(V\) and a functional common-information quantity \(K_{f(X_1,X_2)}\) that can exceed ordinary Gács–Körner–Witsenhausen common information [2103.16717]. The resulting schemes yield achievable rate expressions for computing functions and, in some cases, the sources themselves. The paper is explicit that it does not provide a unified capacity theorem for arbitrary distributed function computation, but it identifies a low-complexity operational principle: function-compression gains arise from task-aligned common structure rather than source correlation alone [2103.16717].

## 4. Task-oriented, goal-dependent, and cognitive formulations

In task-oriented compression, the preserved object is not a source signal but the quality of a downstream optimization or control objective. “Goal-oriented compression for \(L_p\)-norm-type goal functions” studies a source parameter vector \(\ell\in\mathbb R_+^N\) that is used at the receiver to solve
\[
\max_x u(x;\ell),\qquad
u(x;\ell)=-\|x+\ell\|_p,
\]
subject to \(\sum_{j=1}^N x_j\ge E\) and \(x_j\ge 0\) [2405.07808]. Compression is evaluated by
\[
\Gamma=
\mathbb E_\ell\!\left[
\left|
u(x^\star(\ell);\ell)-u(x^\star(\widehat\ell);\ell)
\right|^2
\right],
\]
where \(\widehat\ell=h(q(g(\ell)))\) is produced by precoding, quantization, and decoding. The paper derives a water-filling form for the optimal decision,
\[
x_j^\star=(\mu-\ell_j)^+,
\]
develops linear and nonlinear task-aware transforms, and designs a goal-oriented vector quantizer whose cells minimize excess task loss rather than reconstruction error [2405.07808]. This formulation treats function-compression capacity as a rate–task-loss tradeoff rather than a rate–distortion tradeoff.

A psychologically distinct but structurally related use appears in “Reward function compression facilitates goal-dependent reinforcement learning” [2509.06810]. There the compressed object is the outcome-to-reward mapping itself: humans are instructed that some novel outcome is “Goal” and another is “Nongoal,” and with repeated experience they infer a compressed reward function that discards irrelevant outcome detail while preserving the information needed to assign reward. The paper is explicit that what gets compressed is not the action policy and not the learned stimulus-action values, but the reward function in the standard reinforcement-learning sense, a mapping from outcomes to scalar reward [2509.06810]. Function-compression capacity is not given a formal information-theoretic definition; instead it is operationalized behaviorally through working-memory load, goal-space size, rule complexity, and structural compressibility. Larger across-trial goal spaces impair learning, compressible goal spaces improve it, and faster reward-collection reaction times correlate with higher effective learning rates [2509.06810]. A plausible implication is that this literature treats function-compression capacity as an effective-complexity limit on human valuation, not as a symbolic or Shannon-theoretic bound.

## 5. Low-rank, basis, and wave-function compression

In operator-compression work, capacity is tied to singular spectra and effective rank. “Analytic Origin of Green-Function Compression in the Intermediate Representation” studies finite-temperature imaginary-time Green functions generated by kernels \(K^\alpha(\tau,\omega)\) and defines the effective rank
\[
N_{\rm eff}(\Lambda;\epsilon)=\#\{\,l\mid s_l/s_0>\epsilon\,\},
\]
where \(s_l(\Lambda)\) are singular values of the intermediate-representation kernel [2605.24814]. The paper shows that the fermionic kernel hides a weighted finite-Laplace-transform structure, with a weight-free core diagonalized by oblate spheroidal wave functions. This leads to the asymptotic law
\[
N_{\rm eff}^{F}(\Lambda;\epsilon)\sim
\frac{2}{\pi^2}\,\operatorname{arcosh}(\epsilon^{-2})\,\log\Lambda
\]
for fermions, while bosonic effective rank saturates,
\[
N_{\rm eff}^{B}(\Lambda;\epsilon)\to O(1),
\]
in the low-temperature limit [2605.24814]. Here function-compression capacity is a concrete low-rank approximation property of a kernel-generated function class.

A closely related sensor-limited formulation appears in “Basis function compression for field probe monitoring” [2410.00754]. The monitored magnetic field perturbation is expanded as
\[
\Phi_{\text{calib}}=P_{\text{calib}}\,k_{\text{calib}},
\]
and higher-order basis functions are compressed by principal-component analysis of calibration coefficients. With weighting matrix \(T\) and compression matrix \(C\), the compressed coefficients and compressed probing matrix are
\[
\hat k=C^{T}Tk,\qquad
\tilde P=PT^{-1}C.
\]
The retained latent dimension is \(4+L\), because \(0\)th- and \(1\)st-order terms are preserved explicitly and only \(2\)nd-and-higher orders are compressed [2410.00754]. In the main Western 7T experiment the best performance used \(L=5\), so the total fitted dimension was \(9\), “equivalent to the total number of terms included in a 2nd order fit,” yet the compressed \(5\)th-order fit performed markedly better than conventional fitting [2410.00754]. This is an operational capacity statement: a 9-dimensional learned basis retained more approximation power than a conventional basis of the same dimension.

“Efficient and Scalable Wave Function Compression Using Corner Hierarchical Matrices” addresses a different function class: CASCI/FCI coefficient arrays \(C_{\alpha\beta}\) viewed as a matrix \(\mathbf C\) [2407.21134]. The proposed CHACI method recursively partitions the matrix, emphasizes the upper-left corner after row and column sorting by \(L^2\) norm, and allocates block ranks by the information-density criterion
\[
\rho_i=\frac{\sigma_i^2}{n_{\rm row}+n_{\rm col}+1}.
\]
For the \(14\!-\!14\) active space, CHACI kept singlet-triplet gap error below \(0.07\) eV with only \(28\) kdoubles stored, compared with \(11{,}778\) kdoubles for dense storage, whereas truncated global SVD still had errors around \(0.2\) eV even with \(220\) kdoubles [2407.21134]. For the \(16\!-\!16\) active space, CHACI achieved errors at or below \(0.1\) eV with only \(59\) kdoubles, compared with \(165{,}637\) kdoubles dense storage [2407.21134]. The paper further reports that the compression ratio improves with increasing active-space size and that near-optimal CHACI typically uses about \(2\%\) of the storage of a global TSVD of equal accuracy [2407.21134]. In this domain, function-compression capacity is the retained physical fidelity per stored degree of freedom.

## 6. Learning-theoretic, architectural, and boundary cases

Several papers reinterpret capacity through compression size or structural signal transformation rather than communication cost. “Compression, Generalization and Learning” defines a compression function \(c(U)\subseteq U\) on finite multisets and studies the probability of change of compression,
\[
\Phi_N=
\mathbb P\!\left\{
c(c(\mathbf z_1,\ldots,\mathbf z_N),\mathbf z_{N+1})
\neq
c(\mathbf z_1,\ldots,\mathbf z_N)
\mid \mathbf z_1,\ldots,\mathbf z_N
\right\},
\]
under the preference property [2301.12767]. If \(K=|c(\mathbf z_1,\ldots,\mathbf z_N)|\), then \(K\) yields finite-sample confidence bounds on \(\Phi_N\), and under stronger conditions \(K/N\) is a strongly consistent estimator of \(\Phi_N\) [2301.12767]. When compression change corresponds to prediction error, the observed compression cardinality becomes an operational effective-capacity statistic for learning.

A different structural usage appears in “Neural Network Layer Algebra: A Framework to Measure Capacity and Compression in Deep Learning” [2107.01081]. The paper distinguishes capacity, associated with expressivity, from compression, associated with learnability, using two topology-dependent, parameter-independent metrics: layer complexity \(c^l\) and layer intrinsic power \(p^l\). For kernel or weighting operations, intrinsic power is
\[
p^l=\frac{\Phi_{\rm out}}{\Phi_{\rm in}},
\]
while complexity is
\[
c^l=\log_2(\Phi_{\rm local}).
\]
These local quantities propagate through “layer algebra” to global cumulative intrinsic power (GCIP) and global cumulative complexity (GCC) [2107.01081]. The paper’s interpretation is that high capacity alone is not enough: architectures with similar GCC can differ substantially in GCIP, which it uses to explain why residual networks are easier to train than plain networks of similar expressivity [2107.01081]. This is not source coding, but it is a structural theory of function capacity moderated by compression.

Across the surveyed literature, several open problems recur. The reward-learning work does not formalize compression capacity in bits, dimensions, or rule length and does not model the latent process by which a compressed reward function is inferred [2509.06810]. Helper-based distributed function compression states explicitly that a unified theory of the fundamental limits of functional compression is still lacking [2103.16717]. The vector-linear source-coding paper shows that its general lower bound is not always tight [2508.02996], while the network computing paper proves that its improved upper bound is not generally achievable [1710.02252]. Basis compression for MRI field monitoring depends strongly on the calibration set, and CHACI wave-function compression leaves open the problem of finding an a priori ordering for direct compressed solvers [2410.00754, 2407.21134]. Taken together, these results suggest that function-compression capacity is not a single theory but a family of operational limits: zero-error distinguishability in distributed computation, graph-entropic irreducibility of function-relevant labels, task-preserving compression for optimization and cognition, and effective rank for structured function classes.

Source: https://www.emergentmind.com/topics/function-compression-capacity