Papers
Topics
Authors
Recent
Search
2000 character limit reached

Random Coarse-Graining Framework

Updated 8 July 2026
  • Random coarse-graining is a probabilistic framework that replaces deterministic mapping with a stochastic coarse-to-fine reduction process.
  • It employs latent-variable formulations where coarse variables are random generators, preserving key physical structures and capturing epistemic uncertainty.
  • The framework is utilized across stochastic PDEs, statistical mechanics, and thermodynamics to enable efficient simulation and reliable uncertainty quantification.

Random coarse-graining framework denotes a class of coarse-graining formulations in which the coarse description is itself treated as a random object rather than as the output of a deterministic many-to-one reduction. In the stochastic-PDE setting, the coarse parameters are modeled as random variables conditioned on fine-scale randomness (Grigo et al., 2017). In equilibrium statistical mechanics, coarse variables are treated as latent generators of fine-scale configurations through a probabilistic coarse-to-fine map (Schöberl et al., 2016). In stochastic thermodynamics, the same phrase is used for observation schemes in which the coarse-grained trajectory is not uniquely assigned to a microscopic one, but is produced with a conditional path weight that includes a small chance of error (Meer et al., 3 Jul 2025). Across these formulations, the common objective is to reduce dimensionality while preserving physically relevant structure and quantifying the uncertainty introduced by information loss.

1. Conceptual definition and scope

A random coarse-graining framework replaces deterministic restriction by a probabilistic map between fine and coarse levels. In the predictive coarse-graining formulation for equilibrium ensembles, the central modeling choice is a directed probabilistic model

pˉ(X,x)=pcf(xX)pc(X),\bar p(X,x) = p_{cf}(x \mid X)\, p_c(X),

where the coarse variables XX are latent random variables and the fine configuration xx is sampled from a probabilistic coarse-to-fine map pcf(xX)p_{cf}(x\mid X) (Schöberl et al., 2016). This differs from traditional coarse-graining based on a deterministic fine-to-coarse map X=R(x)X=R(x), for which reconstruction is the uniform distribution over all fine states consistent with XX (Schöberl et al., 2016).

In the stochastic-PDE setting, the same conceptual move appears as a probabilistic coarse-graining operator acting on random media. The fine-scale coefficients λf\lambda_f are high-dimensional random fields, while the effective coarse properties λc\lambda_c are random variables with encoder distribution pc(λcλf,θc)p_c(\lambda_c\mid \lambda_f,\theta_c), rather than deterministic homogenized values (Grigo et al., 2017). The paper explicitly states that this is “precisely a Random Coarse-Graining Framework” because the coarse description is random, conditioned on fine-scale randomness, and Bayesian inference produces random effective parameters representing epistemic uncertainty (Grigo et al., 2017).

A related extension appears in stochastic thermodynamics, where observation itself is treated as probabilistic. Instead of a deterministic map γΓ\gamma\mapsto\Gamma from microscopic trajectory to coarse trajectory, one introduces a conditional path weight XX0, so that the observed coarse-grained trajectory is random even when the microscopic trajectory is fixed (Meer et al., 3 Jul 2025). This shifts the framework from deterministic coarse-graining to random, error-prone observation while retaining lower bounds on entropy production under a time-reversal symmetry condition on the observation process (Meer et al., 3 Jul 2025).

These formulations suggest a broad interpretation: random coarse-graining is not a single algorithm but a family of probabilistic reductions in which latent coarse states, effective parameters, or observations are assigned distributions rather than unique values.

2. Probabilistic latent-variable structure

The most explicit latent-variable formulation is the two-component generative model introduced for coarse-graining in equilibrium statistical mechanics and reused, with PDE-specific structure, in stochastic PDEs. In predictive coarse-graining, the fine-scale marginal induced by the model is

XX1

and learning proceeds by maximizing the log-likelihood of fine-scale samples, or its MAP counterpart with priors on the parameters (Schöberl et al., 2016). The information-theoretic objective is the KL divergence XX2, so maximizing the data log-likelihood is equivalent to minimizing discrepancy between the fine-scale distribution and the distribution induced by the coarse generative model (Schöberl et al., 2016).

For stochastic PDEs, the approximate conditional density is written as

XX3

which decomposes the surrogate into a low-dimensional encoding XX4, a physical coarse solver XX5, and a statistical decoder XX6 (Grigo et al., 2017). The encoder models the log effective conductivity XX7 as a Gaussian linear model over feature functions of the microstructure,

XX8

while the decoder is Gaussian around coarse interpolation,

XX9

with xx0 (Grigo et al., 2017).

The same latent-state architecture appears in data-driven discovery of coarse dynamics, where the coarse-graining process is cast as a probabilistic state-space model. There, the transition law

xx1

governs the CG evolution, and the emission law

xx2

defines the coarse-to-fine map (Felsberger et al., 2018). This formulation treats the CG variables xx3 as hidden random variables and uses Stochastic Variational Inference for posterior inference, with sparse Bayesian learning to identify the salient terms in the CG evolution law (Felsberger et al., 2018).

Across these cases, the latent-variable formulation serves two purposes. First, it replaces a direct map from fine inputs to fine outputs by a lower-dimensional stochastic representation. Second, it turns information loss into an explicit probabilistic object that can be propagated into predictions.

3. Canonical realization for stochastic PDEs

The stochastic-PDE realization is built around a fine model

xx4

with random material properties xx5, and a coarse model

xx6

posed on a much coarser mesh (Grigo et al., 2017). The sample problem is the two-dimensional stationary heat equation on the unit square with a binary random conductivity field generated from a Gaussian process (Grigo et al., 2017).

The fine discretization produces a random input vector xx7 and fine solution xx8, related through the residual equation xx9 (Grigo et al., 2017). In the heat example, pcf(xX)p_{cf}(x\mid X)0 and pcf(xX)p_{cf}(x\mid X)1, so repeated fine solves are computationally expensive (Grigo et al., 2017). The coarse model uses pcf(xX)p_{cf}(x\mid X)2, pcf(xX)p_{cf}(x\mid X)3, or pcf(xX)p_{cf}(x\mid X)4 elements, so pcf(xX)p_{cf}(x\mid X)5 instead of pcf(xX)p_{cf}(x\mid X)6 (Grigo et al., 2017).

The encoder is feature-based and sparse. For each coarse element pcf(xX)p_{cf}(x\mid X)7,

pcf(xX)p_{cf}(x\mid X)8

with 306 candidate features in the reported example (Grigo et al., 2017). A Laplacian prior

pcf(xX)p_{cf}(x\mid X)9

enforces sparsity and performs automatic feature selection (Grigo et al., 2017). Empirically, most coefficients are driven to zero; active features include maximum convex area of high-conductivity blobs, geometric means of conductivities along straight paths, and the log of self-consistent effective medium approximations (Grigo et al., 2017).

Learning is performed by MAP estimation using an EM-type scheme based on Jensen’s inequality. Auxiliary densities X=R(x)X=R(x)0 define lower bounds

X=R(x)X=R(x)1

on each data log-likelihood term, leading to alternating E- and M-steps (Grigo et al., 2017). The optimal X=R(x)X=R(x)2 is the posterior over latent coarse parameters,

X=R(x)X=R(x)3

and some parameter updates are available in closed form, including the encoder and decoder variances (Grigo et al., 2017).

Once trained, prediction proceeds by sampling effective properties X=R(x)X=R(x)4, solving the CG PDE for X=R(x)X=R(x)5, and sampling fine predictions from

X=R(x)X=R(x)6

(Grigo et al., 2017). The posterior predictive distribution yields mean fields and pointwise variances via Monte Carlo. In the heat-equation example, predictive intervals X=R(x)X=R(x)7 around the mean almost always contain the true fine solution across the domain, and scalar outputs such as the temperature at the lower-right corner are tightly centered around the true value with nonzero spread (Grigo et al., 2017).

4. Statistical mechanics and molecular realizations

In equilibrium statistical mechanics, predictive coarse-graining formulates the coarse variables as latent generators of all-atom data and interprets the resulting model as an extension of the relative entropy method (Schöberl et al., 2016). The coarse density and coarse-to-fine map are written in exponential-family form,

X=R(x)X=R(x)8

X=R(x)X=R(x)9

and a hierarchical ARD prior is used to drive many coarse parameters to zero (Schöberl et al., 2016). The paper emphasizes predictive posterior distributions for fine reconstructions and macroscopic observables, not just point estimates, and uses an MC-EM scheme with MCMC in the E-step and Robbins–Monro updates in the M-step (Schöberl et al., 2016).

Two illustrative realizations are provided. For a one-dimensional Ising model, the probabilistic coarse-to-fine map makes each fine spin match its parent coarse spin with probability XX0, and the learned coarse potential contains first-, second-, and third-order interactions (Schöberl et al., 2016). Posterior means and credible intervals for magnetization and correlations improve with more data, while increasing the coarse-graining ratio enlarges predictive uncertainty (Schöberl et al., 2016). For SPC/E water, the coarse variables are molecular centers of mass and the coarse-to-fine map is Gaussian,

XX1

with a coarse potential composed of a fixed Stillinger–Weber term plus a learnable two-body correction expanded in localized basis functions (Schöberl et al., 2016).

A different molecular implementation is spectral matching, which targets kinetic consistency rather than only thermodynamic matching. There the coarse generator XX2 is chosen so that its leading eigenvalue equations approximate those of the fine generator, because rare-event kinetics and slow dynamics are encoded in the low-lying spectrum (Nüske et al., 2019). Depending on the parameterization, spectral matching becomes a linear regression problem for an effective potential or a quadratic programming problem for a position-dependent diffusion (Nüske et al., 2019). This preserves implied timescales and metastable structure more directly than force matching or relative entropy minimization alone (Nüske et al., 2019).

More recent work on coarse-grained Boltzmann generators extends the random coarse-graining viewpoint to flow-based generative models in coarse coordinate space. A coarse-graining map XX3 defines CG coordinates XX4, a neural PMF XX5 is learned by force matching, and a continuous normalizing flow XX6 generates CG samples that are reweighted by

XX7

to recover unbiased CG equilibrium statistics (Chen et al., 11 Feb 2026). This framework combines reduced dimensionality, explicit likelihoods, and exact importance reweighting, and the PMF can be trained from biased atomistic data because the fiber distribution XX8 is unchanged by biases that depend only on XX9 (Chen et al., 11 Feb 2026).

Another line of work learns the coarse representation itself. In invertible coarse-graining, a GNN-based map λf\lambda_f0 and a coarse potential λf\lambda_f1 are optimized jointly, while a decoder and a normalizing flow define a conditional generative model λf\lambda_f2 for back-mapping (Chennakesavalu et al., 2022). Thermodynamic consistency is formulated in a weak sense: coarse sampling from λf\lambda_f3, followed by inversion and reweighting, should reproduce fine-grained averages for a chosen observable class λf\lambda_f4 (Chennakesavalu et al., 2022). The reported applications recover the first two moments of several observables in alanine dipeptide and chignolin, including observables not directly definable in the coarse space (Chennakesavalu et al., 2022).

5. Thermodynamic, nonequilibrium, and error-prone formulations

In nonequilibrium statistical mechanics, coarse-graining is formulated through fluctuations of coarse variables rather than through static latent representations. For general Markov processes, the generalized fluctuation–dissipation relation connects the cumulant generating function of coarse-variable increments to a dissipation potential λf\lambda_f5, which in turn generates the deterministic constitutive law

λf\lambda_f6

(Montefusco et al., 2020). The framework applies to jump processes as well as diffusions, and the paper shows for the reaction λf\lambda_f7 that no Green–Kubo-type diffusion scheme can simultaneously recover the correct entropy, friction matrix, and macroscopic evolution when the coarse fluctuations are genuinely jump-like (Montefusco et al., 2020).

A distinct but related development considers faulty or erroneous coarse-graining. There, observation is modeled by a conditional path weight λf\lambda_f8, and if the observation process obeys the time-reversal symmetry

λf\lambda_f9

then the observed entropy production remains a lower bound on the microscopic entropy production (Meer et al., 3 Jul 2025). For Markov networks with transition misidentification, the paper derives bounds on the sensitivity of the observed entropy production to small errors and shows that redundancies in coarse-grained trajectories can be used to detect and correct some errors (Meer et al., 3 Jul 2025).

The counting-based irreversibility framework makes the random mapping explicit. Instead of deterministic lumping, each fine-grained state λc\lambda_c0 maps probabilistically to coarse state λc\lambda_c1 with λc\lambda_c2 and λc\lambda_c3, giving

λc\lambda_c4

(Bao et al., 15 Aug 2025). Choosing λc\lambda_c5 to be normalized particle number densities in virtual boxes produces a practically simple implementation: irreversibility is inferred from asymmetries of density cross-correlations between boxes, without single-particle tracking or prior knowledge of interactions (Bao et al., 15 Aug 2025). The resulting estimator yields rigorous lower bounds on entropy production and remains applicable to heterogeneous, many-body, periodically driven systems (Bao et al., 15 Aug 2025).

These nonequilibrium and thermodynamic formulations broaden the meaning of random coarse-graining. The random object need not be a latent coarse state alone; it may also be a stochastic observation channel or a random map from fine states to coarse observables.

6. Learning, geometry, and broader extensions

Several recent works extend random coarse-graining beyond classical latent-state models by emphasizing geometry, representation learning, and structured observables. In manifold learning for atomistic-to-continuum coarse-graining, high-dimensional fields from molecular dynamics and continuum PDE models are projected onto a Grassmann manifold, and manifold distances between subspaces are used as the objective for a Gaussian-process surrogate and Efficient Global Optimization (Kontolati et al., 2021). The framework is probabilistic because both continuum parameters and coarse-graining parameters are treated as random variables over prescribed ranges, the misfit is modeled by a Gaussian process, and expected improvement uses the surrogate mean and variance to guide exploration (Kontolati et al., 2021).

RG-motivated learning provides a different notion of coarse random data reduction. It coarse-grains pairs of features chosen to maximize within-pair correlation while minimizing global projection error, leading to a multiscale linear map

λc\lambda_c6

and, in a nonlinear extension, to an information-bottleneck-like objective

λc\lambda_c7

(Landy et al., 2023). The method is tested on random Gaussian data, the Ising model, and glass systems, where it reveals collective modes through COD, MSE, mean-variance scaling, and activity observables (Landy et al., 2023). This suggests a broader use of “random coarse-graining framework” for data-driven reduction in which randomness refers to random or weakly correlated data rather than to latent-variable uncertainty alone.

In networked and quantum settings, structurally constrained coarse-graining plays an analogous role. Spectral coarse graining for bipartite networks groups nodes of the same type by similarity of components in eigenvectors of the two-step random-walk matrices λc\lambda_c8 and λc\lambda_c9, preserving most relevant spectral properties and closely matching mean first passage times after reduction (Wang et al., 2012). In pc(λcλf,θc)p_c(\lambda_c\mid \lambda_f,\theta_c)0-qubit phase space, coarse-graining is built from coset decompositions of finite fields, yielding coarse Wigner functions and systematic subsets of measurement operators; the non-uniqueness of the coset structure makes randomized variants a plausible extension (Matteo et al., 2017).

A plausible implication is that the central design principles are now stable across domains. One chooses coarse variables or observables that retain physically relevant structure, represents the missing information probabilistically or geometrically, and evaluates coarse models by predictive distributions, spectral properties, thermodynamic bounds, or observable preservation. Within that common template, random coarse-graining framework names a shift in emphasis: coarse-graining is treated not merely as compression, but as inference under uncertainty, with the information discarded by reduction made explicit rather than suppressed.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Random Coarse-Graining Framework.