Papers
Topics
Authors
Recent
Search
2000 character limit reached

Recursively Computable Approximate Sufficient Statistics

Updated 17 July 2026
  • RCASS are low-dimensional, recursively updated summaries that approximate sufficient statistics to compress data while retaining essential inferential information.
  • They are applied across diverse domains including approximate Bayesian computation, online learning, and gravitational-wave searches to streamline complex analyses.
  • RCASS enable efficient, scalable inference by reducing storage and computation through recursive update rules and information-theoretic selection criteria.

Recursively Computable Approximate Sufficient Statistics (RCASS) are data summaries that render an approximate likelihood, posterior, regret bound, or state-evolution description a fixed function of a finite-dimensional state, while that state can be updated recursively as new observations, segments, or iterations arrive. Across the cited literature, RCASS appear as information-theoretically selected summaries in approximate Bayesian computation (ABC), polynomial approximate sufficient statistics for generalized linear models, additive sufficient statistics for Burkholder-based online learning, cross-correlation integrands for stochastic gravitational-wave background searches, code-length-optimal approximate sufficient statistics for parametric families, and scalar state variables for sufficient statistic memory approximate message passing (Barnes et al., 2011, Huggins et al., 2017, Foster et al., 2018, Matas et al., 2020, Hayashi et al., 2016, Liu et al., 2022).

1. Definition and formal structure

At the most classical level, sufficiency means that a statistic SS retains all information in the data XX about a parameter θ\theta. The Fisher–Neyman factorization expresses this as

p(xθ)=h(x)g(S(x),θ),p(x\mid \theta)=h(x)\,g(S(x),\theta),

and an equivalent posterior formulation is

p(θx)=p(θS(x)).p(\theta\mid x)=p(\theta\mid S(x)).

In information-theoretic form, sufficiency is equivalent to

I(θ;S)=I(θ;X),I(\theta;S)=I(\theta;X),

or, equivalently, I(θ;XS)=0I(\theta;X\mid S)=0 (Barnes et al., 2011).

Approximate sufficiency relaxes this equality by minimizing the information lost when S(X)S(X) replaces XX. In the ABC setting, the loss for parameter inference is

Lparam(S)=EX ⁣[DKL ⁣(p(θX)p(θS(X)))],L_{\mathrm{param}}(S)=\mathbb{E}_{X}\!\left[D_{\mathrm{KL}}\!\big(p(\theta\mid X)\,\|\,p(\theta\mid S(X))\big)\right],

and for model selection

XX0

A statistic is approximately sufficient when the relevant loss is small (Barnes et al., 2011).

The “recursively computable” part refers to update rules that avoid storing the full history. In additive settings this takes the form

XX1

while in more general recursive settings one writes

XX2

The central operational idea is that both inference and computation are carried by the evolving state XX3, not by the raw data sequence (Foster et al., 2018).

This suggests that RCASS is best understood as a cross-domain design principle: one searches for a low-dimensional state that is approximately sufficient for the target task and closed under streaming, distributed, or iterative updates.

Setting RCASS state Recursive mechanism
ABC selected summary subset add statistics that most reduce KL loss
PASS-GLM XX4 or XX5 additive updates
Online learning XX6 additive updates
SGWB searches frequency integrand and inverse variance segmentwise weighted sums
Parametric coding quantized estimator or expectation-parameter index update estimate, then re-quantize
SS-MAMP XX7 state-evolution recursion

2. Information-theoretic RCASS in approximate Bayesian computation

In ABC, likelihoods are intractable but simulation from the model is available. The posterior based on a summary statistic XX8 is

XX9

with the paper focusing on the special case θ\theta0. The methodological problem is that comparing summaries rather than full data generally loses information unless the summaries are sufficient (Barnes et al., 2011).

The RCASS construction in this setting is explicitly recursive. Starting from a candidate pool θ\theta1, one either performs an exhaustive search over all subsets or, more practically, a greedy forward construction. The greedy parameter-inference procedure initializes θ\theta2, chooses the first statistic to maximize mutual information with θ\theta3, and thereafter adds the statistic that maximizes the KL change

θ\theta4

The recursion stops when the marginal gain falls below a tolerance θ\theta5, and an optional backward elimination stage re-tests previously selected summaries for redundancy. The paper also notes an optional penalized objective,

θ\theta6

to favor parsimony (Barnes et al., 2011).

For model selection, the construction is more restrictive because parameter sufficiency within each model does not by itself imply sufficiency for discriminating models. The joint-space criterion is

θ\theta7

with decomposition

θ\theta8

The practical consequence is a two-stage RCASS workflow: first construct minimal sufficient sets for parameters within each model, then add further statistics that materially change the model posterior or Bayes factors conditional on those parameter summaries (Barnes et al., 2011).

The empirical demonstrations show why this distinction matters. In the two-normal example, the mean θ\theta9 is selected universally for parameter inference, while model selection requires adding p(xθ)=h(x)g(S(x),θ),p(x\mid \theta)=h(x)\,g(S(x),\theta),0, so that p(xθ)=h(x)g(S(x),θ),p(x\mid \theta)=h(x)\,g(S(x),\theta),1 is sufficient for the joint space. In the coalescent-model experiments, selected summaries vary across datasets and models, with p(xθ)=h(x)g(S(x),θ),p(x\mid \theta)=h(x)\,g(S(x),\theta),2 frequently selected and the random statistic p(xθ)=h(x)g(S(x),θ),p(x\mid \theta)=h(x)\,g(S(x),\theta),3 rarely chosen. In the random-walk models, p(xθ)=h(x)g(S(x),θ),p(x\mid \theta)=h(x)\,g(S(x),\theta),4 is often selected across models, while p(xθ)=h(x)g(S(x),θ),p(x\mid \theta)=h(x)\,g(S(x),\theta),5 is preferred particularly for biased walks. These cases exemplify data-dependent approximate sufficiency rather than a fixed universal summary set (Barnes et al., 2011).

The framework also makes the main limitations explicit. Greedy selection is order-dependent, high-dimensional summary spaces induce the ABC curse of dimensionality, multivariate KL estimation can be noisy, and summaries sufficient for parameter inference may remain insufficient for model choice. The prescribed mitigations are backward elimination, stochastic or heuristic search, scale-aware discrepancies, sensitivity analysis over p(xθ)=h(x)g(S(x),θ),p(x\mid \theta)=h(x)\,g(S(x),\theta),6, and explicit enforcement of joint-space sufficiency (Barnes et al., 2011).

3. Polynomial and distributed RCASS for Bayesian generalized linear models

PASS-GLM instantiates RCASS by replacing the non-linear term in a generalized linear model log-likelihood with a polynomial approximation. For canonical GLMs,

p(xθ)=h(x)g(S(x),θ),p(x\mid \theta)=h(x)\,g(S(x),\theta),7

and PASS-GLM replaces p(xθ)=h(x)g(S(x),θ),p(x\mid \theta)=h(x)\,g(S(x),\theta),8 by

p(xθ)=h(x)g(S(x),θ),p(x\mid \theta)=h(x)\,g(S(x),\theta),9

Because

p(θx)=p(θS(x)).p(\theta\mid x)=p(\theta\mid S(x)).0

the approximate log-likelihood depends on the data only through sums of tensor powers of the covariates, yielding approximate sufficient statistics that are additive across data points, updated recursively in streaming, aggregated exactly by summation in distributed settings, and introduce no additional approximation at aggregation time (Huggins et al., 2017).

In the logistic parameterization with p(θx)=p(θS(x)).p(\theta\mid x)=p(\theta\mid S(x)).1, the paper uses a polynomial approximation

p(θx)=p(θS(x)).p(\theta\mid x)=p(\theta\mid S(x)).2

in an orthogonal basis, specifically Chebyshev polynomials on p(θx)=p(θS(x)).p(\theta\mid x)=p(\theta\mid S(x)).3. The coefficients satisfy

p(θx)=p(θS(x)).p(\theta\mid x)=p(\theta\mid S(x)).4

and the monomial coefficients are obtained through the basis expansion p(θx)=p(θS(x)).p(\theta\mid x)=p(\theta\mid S(x)).5. The approximate likelihood then takes the exponential-family-like form

p(θx)=p(θS(x)).p(\theta\mid x)=p(\theta\mid S(x)).6

where

p(θx)=p(θS(x)).p(\theta\mid x)=p(\theta\mid S(x)).7

These p(θx)=p(θS(x)).p(\theta\mid x)=p(\theta\mid S(x)).8 are the RCASS in the precise sense that the whole approximation is a fixed function of p(θx)=p(θS(x)).p(\theta\mid x)=p(\theta\mid S(x)).9 and the accumulated statistics (Huggins et al., 2017).

The recursive update rule is purely additive. When a new datum I(θ;S)=I(θ;X),I(\theta;S)=I(\theta;X),0 arrives,

I(θ;S)=I(θ;X),I(\theta;S)=I(\theta;X),1

For quadratic logistic regression, the RCASS reduce to

I(θ;S)=I(θ;X),I(\theta;S)=I(\theta;X),2

with updates

I(θ;S)=I(θ;X),I(\theta;S)=I(\theta;X),3

With a Gaussian prior and I(θ;S)=I(θ;X),I(\theta;S)=I(\theta;X),4, the approximate posterior is Gaussian with closed-form mean and covariance (Huggins et al., 2017).

Theoretical guarantees are attached to the polynomial approximation itself. Chebyshev approximation error on I(θ;S)=I(θ;X),I(\theta;S)=I(\theta;X),5 is exponentially small in the degree I(θ;S)=I(θ;X),I(\theta;S)=I(\theta;X),6, and the paper gives bounds for MAP error, approximate posterior quality in Wasserstein-I(θ;S)=I(θ;X),I(\theta;S)=I(\theta;X),7 distance, and posterior mean and uncertainty estimates. For logistic regression, the reported approximation obeys

I(θ;S)=I(θ;X),I(\theta;S)=I(\theta;X),8

with the empirical observation that for many datasets at least I(θ;S)=I(θ;X),I(\theta;S)=I(\theta;X),9 of the inner products I(θ;XS)=0I(\theta;X\mid S)=00 fall in I(θ;XS)=0I(\theta;X\mid S)=01. The paper also states a pathology for degrees I(θ;XS)=0I(\theta;X\mid S)=02, I(θ;XS)=0I(\theta;X\mid S)=03, and recommends I(θ;XS)=0I(\theta;X\mid S)=04 (Huggins et al., 2017).

The empirical results show how these structural claims translate into computation. PASS-LR2 on ChemReact, Webspam, CovType, and CodRNA was reported as approximately I(θ;XS)=0I(\theta;X\mid S)=05 faster than SGD and I(θ;XS)=0I(\theta;X\mid S)=06–I(θ;XS)=0I(\theta;X\mid S)=07 faster than Laplace, while remaining competitive in posterior mean, posterior variance, and test log-likelihood; CodRNA is identified as a known outlier because many I(θ;XS)=0I(\theta;X\mid S)=08 fall outside I(θ;XS)=0I(\theta;X\mid S)=09. On Criteo, with S(X)S(X)0 million points and S(X)S(X)1 dimensions after random projection, streaming PASS-LR2 achieved negative test log-likelihood S(X)S(X)2 versus S(X)S(X)3 for SGD, slightly worse AUC, and distributed near-linear speedups with S(X)S(X)4 cores giving approximately S(X)S(X)5 speedup from a baseline of S(X)S(X)6 minutes at S(X)S(X)7 core on S(X)S(X)8 million points (Huggins et al., 2017).

4. Additive RCASS in online learning and the Burkholder method

In online learning, RCASS arise when regret can be upper bounded by a function of cumulative sufficient statistics rather than the full data sequence. The paper defines a sufficient statistic pair S(X)S(X)9 by the inequality

XX0

This representation is “approximate” when the desired regret expression is upper bounded by XX1 rather than exactly equal to it. The recursively computable state is

XX2

so the learner need only keep XX3 in memory (Foster et al., 2018).

The Burkholder method gives the structural condition under which such RCASS support an actual online algorithm. The key object is a function XX4 satisfying XX5, XX6, and the restricted concavity condition

XX7

for mean-zero XX8. In the convex case it suffices to verify a Rademacher version. Existence of such a Burkholder function is equivalent to the needed martingale inequality, and the resulting prediction strategy evaluates XX9 only on the cumulative sufficient statistic (Foster et al., 2018).

Two explicit instantiations make the RCASS structure concrete. For parameter-free supervised learning with linear classes and a Lparam(S)=EX ⁣[DKL ⁣(p(θX)p(θS(X)))],L_{\mathrm{param}}(S)=\mathbb{E}_{X}\!\left[D_{\mathrm{KL}}\!\big(p(\theta\mid X)\,\|\,p(\theta\mid S(X))\big)\right],0-smooth norm, the sufficient statistics are

Lparam(S)=EX ⁣[DKL ⁣(p(θX)p(θS(X)))],L_{\mathrm{param}}(S)=\mathbb{E}_{X}\!\left[D_{\mathrm{KL}}\!\big(p(\theta\mid X)\,\|\,p(\theta\mid S(X))\big)\right],1

so the running state is Lparam(S)=EX ⁣[DKL ⁣(p(θX)p(θS(X)))],L_{\mathrm{param}}(S)=\mathbb{E}_{X}\!\left[D_{\mathrm{KL}}\!\big(p(\theta\mid X)\,\|\,p(\theta\mid S(X))\big)\right],2, maintained in Lparam(S)=EX ⁣[DKL ⁣(p(θX)p(θS(X)))],L_{\mathrm{param}}(S)=\mathbb{E}_{X}\!\left[D_{\mathrm{KL}}\!\big(p(\theta\mid X)\,\|\,p(\theta\mid S(X))\big)\right],3 time and space per step. The corresponding time-varying Burkholder family is

Lparam(S)=EX ⁣[DKL ⁣(p(θX)p(θS(X)))],L_{\mathrm{param}}(S)=\mathbb{E}_{X}\!\left[D_{\mathrm{KL}}\!\big(p(\theta\mid X)\,\|\,p(\theta\mid S(X))\big)\right],4

which yields the comparator-dependent regret bound stated in the paper (Foster et al., 2018).

For matrix prediction, the sufficient statistics are

Lparam(S)=EX ⁣[DKL ⁣(p(θX)p(θS(X)))],L_{\mathrm{param}}(S)=\mathbb{E}_{X}\!\left[D_{\mathrm{KL}}\!\big(p(\theta\mid X)\,\|\,p(\theta\mid S(X))\big)\right],5

with

Lparam(S)=EX ⁣[DKL ⁣(p(θX)p(θS(X)))],L_{\mathrm{param}}(S)=\mathbb{E}_{X}\!\left[D_{\mathrm{KL}}\!\big(p(\theta\mid X)\,\|\,p(\theta\mid S(X))\big)\right],6

The corresponding Burkholder function,

Lparam(S)=EX ⁣[DKL ⁣(p(θX)p(θS(X)))],L_{\mathrm{param}}(S)=\mathbb{E}_{X}\!\left[D_{\mathrm{KL}}\!\big(p(\theta\mid X)\,\|\,p(\theta\mid S(X))\big)\right],7

certifies a variance-adaptive regret bound involving the spectral norm of Lparam(S)=EX ⁣[DKL ⁣(p(θX)p(θS(X)))],L_{\mathrm{param}}(S)=\mathbb{E}_{X}\!\left[D_{\mathrm{KL}}\!\big(p(\theta\mid X)\,\|\,p(\theta\mid S(X))\big)\right],8. The RCASS state is Lparam(S)=EX ⁣[DKL ⁣(p(θX)p(θS(X)))],L_{\mathrm{param}}(S)=\mathbb{E}_{X}\!\left[D_{\mathrm{KL}}\!\big(p(\theta\mid X)\,\|\,p(\theta\mid S(X))\big)\right],9, so storage and computation depend on the maintained aggregates rather than on XX00 (Foster et al., 2018).

A common misconception is that online sufficient statistics in this framework are merely bookkeeping devices. In fact, the paper shows that the existence of the sufficient statistic representation and the existence of a Burkholder function are jointly algorithmic and analytic: the same state both certifies the martingale inequality and drives the online strategy (Foster et al., 2018).

5. Cross-correlation RCASS in stochastic gravitational-wave background searches

For stochastic gravitational-wave background searches with ground-based interferometers, the approximate sufficient statistics are the frequency integrand of the cross-correlation statistic and its variance. The sufficiency is approximate because the analysis uses the weak-signal approximation,

XX01

and replaces the unknown auto-powers by measured estimates from neighboring segments, with a bias-correction factor XX02 (Matas et al., 2020).

At coarse-grained frequency bin XX03, the segment-level cross-power estimator is

XX04

and in the XX05 parameterization the sufficient statistics are

XX06

For a fixed spectral shape XX07, they reduce further to the amplitude-level pair XX08 (Matas et al., 2020).

The reduced likelihood depends on the data only through these statistics. In the XX09 representation,

XX10

This is precisely the factorization-theorem statement of approximate sufficiency under the reduced model (Matas et al., 2020).

The RCASS property is explicit at the segment level. If

XX11

and

XX12

then the per-frequency accumulated statistics are

XX13

and a new segment updates them by simple addition. The same recursive accumulation holds for the amplitude-only summaries XX14 (Matas et al., 2020).

The principal significance of this reduction is methodological equivalence. The paper proves that LIGO–Virgo’s hybrid frequentist–Bayesian analysis—frequentist cross-correlation estimation followed by Bayesian parameter estimation on the integrands and variances—is approximately equivalent to a fully Bayesian analysis on the raw time-series under the stated approximations. The practical gain is extreme data-volume reduction: months-long time-frequency data are compressed to two frequency series, or further to a single amplitude estimate and variance for fixed spectral shape (Matas et al., 2020).

The limits of the approximation are also clear. Using the same analysis segment to estimate auto-powers induces bias; the weak-signal approximation breaks down for strong signals; and the treatment does not incorporate non-stationarity beyond the segment model, calibration uncertainties, or correlated non-gravitational-wave noise such as Schumann resonances (Matas et al., 2020).

6. Rate theory and code-length optimal approximate sufficient statistics

A different branch of the literature studies approximate sufficient statistics as compressed codes for parametric families. For XX15 i.i.d. samples from a XX16-nomial family with XX17 degrees of freedom, storing the exact sufficient statistic requires code length

XX18

because the number of types is XX19. The main result is that if small approximation error is allowed, the code length can be reduced to

XX20

and this rate is both achievable and, under the stated criteria, unavoidable (Hayashi et al., 2016).

The paper distinguishes blind and visible encoders and studies errors under KL and TV criteria. Its summarized rate statements include

XX21

XX22

and, under exponential-family structure,

XX23

The paper also proves strong converses: even if non-vanishing error is allowed, the pre-log factor cannot be reduced below XX24 (Hayashi et al., 2016).

The information-theoretic mechanism behind these rates is local Gaussian geometry. Clarke–Barron asymptotics for the Bayesian mixture give

XX25

while local asymptotic normality yields

XX26

Quantization at resolution XX27 therefore leads naturally to a grid of size XX28 and hence to code length XX29 (Hayashi et al., 2016).

The RCASS viewpoint enters through recursive maintenance of low-dimensional estimators or sufficient statistics followed by periodic quantization. In exponential families one may maintain

XX30

quantize the expectation parameter on a lattice of span proportional to XX31, and decode using a conditional mixture over sufficient-statistic fibers. In more general parametric families, one may instead maintain a running estimator and quantize it in the Fisher metric. The paper’s synthesis describes geometric or doubling schedules for emitting the quantized index, keeping the cumulative code length at approximately XX32 with XX33 working memory (Hayashi et al., 2016).

A recurrent misunderstanding is that approximate sufficient statistics in this sense merely approximate a parameter. The actual object is stronger: a compressed index from which one reconstructs a full distribution on XX34 with controlled KL or TV error. The distinction is what makes the strong converse results nontrivial (Hayashi et al., 2016).

7. Sufficient-statistic memory AMP and asymptotic RCASS

In approximate message passing for large random linear systems,

XX35

state evolution describes the dynamics but does not by itself ensure convergence. Sufficient Statistic Memory Approximate Message Passing (SS-MAMP) addresses this by imposing a sufficient statistic condition on a memory-AMP iteration and constructing it by damping from an arbitrary MAMP (Liu et al., 2022).

The general MAMP structure is

XX36

with memory linear estimator XX37 and memory nonlinear estimator XX38. Orthogonality conditions,

XX39

play the Onsager-like role that preserves Gaussianity and valid state evolution (Liu et al., 2022).

The sufficient statistic condition is

XX40

Under the paper’s assumptions, this is equivalent to the covariance matrices being L-banded:

XX41

For Gaussian observation vectors, the paper proves that the last component is sufficient if and only if the last row and column of the covariance are constant, which is the same L-banded structure in finite-dimensional form (Liu et al., 2022).

Once L-bandedness holds, the matrix-valued state evolution collapses to a scalar recursion,

XX42

The RCASS is therefore

XX43

a finite-dimensional state that fully determines the asymptotic dynamics and is recursively computable. The diagonal variance sequences are monotonically nonincreasing and convergent, so the RCASS recursion itself converges (Liu et al., 2022).

The construction from an arbitrary MAMP to SS-MAMP uses optimal damping. If XX44 and XX45 are the raw outputs, the damped outputs are

XX46

with

XX47

This damping preserves orthogonality, yields L-banded covariance matrices, and ensures convergence while retaining correct state evolution (Liu et al., 2022).

This setting makes explicit an asymptotic version of RCASS. The sufficient statistic is not a classical finite-sample statistic for a parametric likelihood; it is a low-dimensional state that is sufficient in the state-evolution sense, with all asymptotic empirical laws parameterized by XX48. The paper’s contribution is to show that, under the sufficient statistic condition, memory and damping can be organized so that convergence and state evolution coexist rather than conflict (Liu et al., 2022).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Recursively Computable Approximate Sufficient Statistics (RCASS).