---
title: 'CosmoGen: Cosmological Model Generation'
url: https://www.emergentmind.com/topics/cosmogen
type: topic
---

# CosmoGen: Cosmological Model Generation

CosmoGen most specifically denotes a cosmological model generator that couples evolutionary symbolic regression to Boltzmann and Bayesian inference codes in order to generate interpretable dark-energy models targeted at specific observational goals, notably the alleviation of the $H_0$ and $S_8$ tensions [2509.15453]. In adjacent literature, the same label also appears as an umbrella for cosmological generative pipelines spanning initial-condition construction, mock-halo synthesis, indirect-detection source spectra, and line-intensity-mapping catalog generation, and in CMS it is used colloquially for the cosmic-muon event generator CMSCGEN [2006.01841] [2402.17581] [2312.01153] [2506.16843] [1107.0427].

## 1. Nomenclature and scope

The term “CosmoGen” is not univocal. In its strictest sense, it names the framework introduced in “CosmoGen: a cosmological model generator” [2509.15453], where the objective is automated model discovery in precision cosmology. There, candidate dark-energy fluids are generated symbolically, inserted into a modified cosmological solver, and ranked against data.

A second usage is infrastructural rather than nominal. Several recent works describe modules that either serve a “CosmoGen pipeline” or are presented explicitly in the context of cosmological generative modeling: genetIC for constrained and genetically modified zoom initial conditions [2006.01841], CosmoMIA for converting coarse-grid halo counts into tracer coordinates and velocities in real and redshift space [2402.17581], CosmiXs for tabulated dark-matter source spectra used in indirect-detection source terms [2312.01153], and CosmoGLINT for Transformer-based generation of galaxy populations from dark-matter-only halo catalogs [2506.16843]. This suggests a broader usage in which “CosmoGen” denotes an end-to-end computational ecosystem for cosmological forward modeling rather than a single executable.

A third usage is domain-specific and unrelated to cosmological model discovery. Within CMS, “CosmoGen” is a colloquial name for CMSCGEN, the dedicated cosmic muon Monte Carlo event generator used in the CMS detector simulation chain [1107.0427]. The coexistence of these meanings is a recurring source of ambiguity.

## 2. Core framework: evolutionary generation of dark-energy models

In the 2025 usage, CosmoGen is a model-building framework that leverages evolutionary symbolic regression tightly coupled to cosmological computations to automatically generate interpretable dark-energy models [2509.15453]. Its proof-of-concept problem is narrowly defined: generate homogeneous dark-energy fluid models that can alleviate the cosmological tensions in $H_0$ and $S_8$ without abandoning a direct confrontation with CMB and weak-lensing information.

The generated models are represented through a normalized dark-energy density,
$$
\Omega_{\mathrm{DE}}(a) = \frac{A}{f(a;D)},
$$
where $a$ is the scale factor, $D$ is a free parameter, and $A$ is fixed by closure. Using
$$
\Omega_{\mathrm{DE},0} = 1 - \sum_i \Omega_i - \Omega_k,
$$
the framework sets
$$
A = f(1;D)\,\Omega_{\mathrm{DE},0},
$$
so that
$$
\Omega_{\mathrm{DE}}(a) = \frac{f(1;D)\,\Omega_{\mathrm{DE},0}}{f(a;D)}.
$$
The symbolic forms for $f(a;D)$ are produced by genetic programming using the grammar with operators $\{\mathrm{Add}, \mathrm{Sub}, \mathrm{Mul}, \mathrm{Div}, \mathrm{Pow}, \mathrm{Inv}, \mathrm{Exp}, \mathrm{Log}, \mathrm{Neg}\}$ acting on terminals $\{a, D, \text{constants}\}$ [2509.15453].

This representation is deliberately constrained. It keeps the search space interpretable, preserves direct control over the background expansion history, and permits immediate insertion into a modified CLASS run. A plausible implication is that the framework is designed less as an unconstrained symbolic-search engine than as a targeted mechanism for constructing viable analytic deformations of $\Lambda$CDM.

## 3. Evolutionary loop, viability filters, and scoring

CosmoGen integrates a genetic-programming engine with CLASS and MontePython so that each candidate symbolic expression is turned into a working cosmological component, tested numerically, and scored against a tension-aware likelihood [2509.15453]. The initialization uses a population size of 4096 random expressions, with a maximum initial complexity of 7 nodes. The evolutionary loop runs for 8 generations, with $\mu=128$ parents, $\lambda=512$ children, mate and mutate probabilities of $0.5$, and a maximum complexity increase per generation of 2.

Before full evaluation, candidates undergo simplification and pre-testing. SymPy and custom routines are used to simplify expressions, transform them into an energy density, test uniqueness, and reject numerical pathologies such as divergences and complex values over representative $D$ values $(0.05, 0.95, 1000)$. Physical viability heuristics require $w(a)$ in the broad range $-15 \leq w \leq 0$ and small adiabatic sound speed $c_s^2 \approx 0$; the best trial values of $(D,A)$ are stored to seed the MCMC [2509.15453].

The candidate is then inserted on-the-fly into a modified CLASS in which $\Lambda$ is replaced by a homogeneous dark-energy fluid with $\Omega_\Lambda=0$. MontePython performs short Metropolis–Hastings chains in the reduced parameter space $(D,h,\omega_{\mathrm{CDM}})$, with all other parameters fixed to Planck 2018 $\Lambda$CDM means during the generation stage. The likelihood combines Planck 2018 high-$\ell$ TTTEEE (lite), low-$\ell$ EE, and low-$\ell$ TT likelihoods with Gaussian priors on SH0ES $H_0$ and KiDS-1000 $S_8$ [2509.15453].

The optimization target is not the raw minimum $\chi^2$ alone. CosmoGen defines
$$
\mathrm{score} = \chi^2_{\min} + C,
$$
where $C$ is the expression-tree complexity. Since typical $\chi^2$ values are $O(10^2$–$10^3)$ and $C$ is $O(1$–$10)$, the complexity penalty acts as a tiebreaker rather than as the dominant regularizer [2509.15453]. Tournament selection then chooses parents for the next generation.

| Component | Setting | Value |
|---|---|---|
| Initial population | Random expressions | 4096 |
| Generations | Evolution limit | 8 |
| Parent/child strategy | $(\mu+\lambda)$ | $\mu=128$, $\lambda=512$ |
| Mate/mutate probabilities | Per offspring step | 0.5 / 0.5 |
| Initial complexity cap | Generation 1 | 7 nodes |
| Complexity growth cap | Per generation | 2 nodes |

The associated cosmological computations are standard but central to the loop. For the implemented flat case,
$$
H^2(a) = H_0^2 \left[\Omega_m a^{-3} + \Omega_r a^{-4} + \Omega_{\mathrm{DE}}(a)\right].
$$
From the continuity equation,
$$
\frac{d\rho_{\mathrm{DE}}}{d\ln a} = -3(1+w)\rho_{\mathrm{DE}},
$$
CosmoGen derives
$$
w(a) = -1 + \frac{1}{3}\frac{d\ln\rho_{\mathrm{DE}}}{d\ln a}.
$$
CLASS then computes $P(k,z)$, $C_\ell$, $\sigma_8$, and derived parameters including
$$
S_8 = \sigma_8 \left(\frac{\Omega_m}{0.3}\right)^{1/2}.
$$
This makes the search intrinsically physics-aware rather than purely symbolic [2509.15453].

## 4. Generated models, the CG example, and empirical results

The pipeline yields $O(10^2)$ viable final models, and the highest-scoring expressions are compact. The illustrative forms reported include
$$
\Omega_{\mathrm{DE}}(a) \propto \frac{1}{Da-\ln a}, \quad
\Omega_{\mathrm{DE}}(a) \propto \frac{1}{Da-\ln(Da)}, \quad
\Omega_{\mathrm{DE}}(a) \propto \frac{1}{D^2a^2-\ln a}, \quad
\Omega_{\mathrm{DE}}(a) \propto \frac{1}{D^3a^3-\ln a}
$$
[2509.15453]. Many of the top candidates therefore combine logarithms with low-degree polynomials in $a$ and $D$, consistent with the framework’s built-in preference for interpretable expressions.

The representative “CG model” selected for deeper analysis has
$$
\rho_{\mathrm{CG}}(a) = \frac{A}{Da-\ln(Da)},
$$
with equation of state
$$
w_{\mathrm{CG}}(a) = -1 + \frac{1}{3}\frac{d\ln\rho}{d\ln a}
= -1 - \frac{aD-1}{3\left(Da-\ln(Da)\right)}.
$$
The reported interpretation is that for $D>0$, $w(a)$ can cross the phantom divide and return, whereas for $D<0$, the evolution is slower and less extreme [2509.15453].

The final Bayesian assessment differs from the tension-aware generation stage. It uses standard datasets without explicit $H_0$ or $S_8$ priors: Planck 2018 and KiDS–VIKING KV-450. The varied parameters are $\log D \in [-5,3]$, $A_s$, $n_s$, $\Omega_b$, $\Omega_{\mathrm{CDM}}$, and $h$, with derived $\Omega_{\mathrm{DE}}$ or $\Omega_\Lambda$, plus nuisance parameters $A_{\mathrm{Planck}}$ and $A_{\mathrm{IA}}$. Spatial flatness is assumed [2509.15453].

| Dataset | CG model | $\Lambda$CDM |
|---|---|---|
| Planck 2018 $H_0$ | $69.4^{+1.3}_{-1.7}$ | $67.88\pm0.62$ |
| Planck 2018 $S_8$ | $0.820^{+0.020}_{-0.018}$ | $0.852\pm0.016$ |
| KV-450 $H_0$ | $74.2^{+7.1}_{-4.3}$ | $73.8^{+6.4}_{-4.6}$ |
| KV-450 $S_8$ | $0.757^{+0.045}_{-0.062}$ | $0.737^{+0.038}_{-0.031}$ |
| Planck–KV-450 $S_8$ tension | $1.1\sigma$ | $2.8\sigma$ |

The reported evidence remains mildly unfavorable to the generated model: $\ln B_{\mathrm{CG},\Lambda}=-1.27$ for Planck 2018 and $\ln B_{\mathrm{CG},\Lambda}=-0.99$ for KV-450, both indicating weak preference for $\Lambda$CDM [2509.15453]. Nonetheless, the framework demonstrates that a machine-generated analytic fluid can shift the Planck-inferred $H_0$ upward by approximately $1\sigma$ and reduce the inter-dataset $S_8$ tension from $2.8\sigma$ to $1.1\sigma$.

A common misconception is that CosmoGen is presented as a wholesale replacement for model comparison. The reported results do not support that interpretation. The Bayesian evidence still weakly prefers $\Lambda$CDM, and the generation-stage likelihood uses targeted priors explicitly intended to encourage tension alleviation [2509.15453]. A more precise characterization is that CosmoGen is a model-discovery engine for proposing interpretable extensions that merit subsequent conventional inference.

## 5. CosmoGen as a broader computational ecosystem

In several neighboring contexts, “CosmoGen” functions as an umbrella for modular cosmological generation workflows rather than for symbolic model discovery alone. The modules described in that literature span initial conditions, halo catalog generation, source spectra, and galaxy-population synthesis.

genetIC provides cosmological initial conditions for $N$-body simulations, including nested zoom regions and “genetic modifications” that alter halo history, mass, or environment while maintaining consistency with the statistical ensemble [2006.01841]. It constructs Gaussian random fields using a white-noise representation,
$$
\delta(k) = T(k)\,n(k), \qquad T(k)\equiv [P(k)]^{1/2},
$$
and computes 1LPT displacements
$$
\psi^{(1)}(k) = -i\,k\,\delta(k)/k^2.
$$
For multi-resolution zooms it introduces complementary Fourier filters,
$$
F_L(k) = \left[\exp\left(\frac{k-k_0}{k_0T}\right)+1\right]^{-1}, \qquad F_L(k)^2+F_H(k)^2=1,
$$
with default choices $k_0 \approx 0.5\,k_{\mathrm{nyq}}$ and $T\approx0.1$ [2006.01841]. The code reports sub-percent precision for modifications within zoom initial conditions and avoids the large ghost-padding memory cost of traditional nested-grid convolution schemes.

CosmoMIA addresses a different stage of the forward model: it takes coarse-grid dark-matter realizations and target number counts per cell and generates coordinates and velocities for individual tracers [2402.17581]. Its four stages are tracer assignment, attractor identification, tracer collapse, and redshift-space distortions. At $z=1.1$, it reproduces AbacusSummit halo clustering with the power-spectrum monopole within $1\%$ up to $k\approx0.6\,h\,\mathrm{Mpc}^{-1}$ in real space, within $1\%$ to $k\approx0.5$ and $2\%$ to $k\approx0.7\,h\,\mathrm{Mpc}^{-1}$ in redshift space, and the quadrupole within $5\%$ to $k\approx0.4\,h\,\mathrm{Mpc}^{-1}$ [2402.17581]. The coherent velocity field is enhanced through
$$
v'_{\mathrm{coh}}(k)=v_{\mathrm{coh}}(k)\,K(k), \qquad
K(k)=1+\frac{\Gamma}{1+e^{-q}k^2},
$$
with fiducial $\Gamma=10$ and $q=0.6$, before mapping to redshift space via
$$
s=x+\frac{(v\cdot\hat n)}{aH}\hat n
$$
[2402.17581].

CosmiXs supplies tabulated source spectra for indirect dark-matter searches and is explicitly described as ready to plug into “CosmoGen” pipelines [2312.01153]. It tabulates $dN/d\log_{10}(x)$ for $\gamma$, $e^+$, $\bar p$, and $\nu_e,\nu_\mu,\nu_\tau$ for 64 dark-matter masses between 5 GeV and 100 TeV and for standard fermionic and bosonic two-body final states, including helicity- and polarization-resolved variants [2312.01153]. The conversion to injection spectra is
$$
\frac{dN}{dE}=\frac{1}{E\ln 10}\frac{dN}{d\log_{10}(x)}, \qquad x\equiv E/m_\chi.
$$
For annihilation and decay, the source terms are
$$
Q_i(E,r)=\frac{1}{2}\langle\sigma v\rangle \frac{\rho(r)^2}{m_\chi^2}\frac{dN_i}{dE}, \qquad
Q_i(E,r)=\frac{\rho(r)}{m_\chi\tau_\chi}\frac{dN_i}{dE},
$$
respectively [2312.01153].

CosmoGLINT extends the generative interpretation of “CosmoGen” to galaxy populations and line-intensity-mapping mocks [2506.16843]. It is a decoder-only Transformer with 4 layers, 8 heads, hidden size 128, and 100-bin softmax heads per property, trained on IllustrisTNG to generate variable-length galaxy sequences conditioned on halo mass. For each halo it emits tuples
$$
x=(\mathrm{SFR}, d, v_r, v_t),
$$
with factorized conditional probability
$$
p(x|h)=\prod_i p(x_i|x_{<i},h)
$$
under causal masking [2506.16843]. At $z=2$, training uses halos with $M>10^{11}M_\odot$ and galaxies with $\mathrm{SFR}>10^{-3}M_\odot/\mathrm{yr}$, yielding 330,287 halos and 5,027,636 galaxies; the reported generation throughput is about 250,000 galaxies/sec or about 16,000 halos/sec [2506.16843]. It reproduces the voxel intensity distribution and the power spectrum in both real and redshift space and is trained across 15 snapshots from $z=0.5$ to $6.0$ for lightcone generation [2506.16843].

Taken together, these modules define a broad computational interpretation of CosmoGen: initial conditions, approximate gravity, tracer realization, galaxy population synthesis, and particle-physics source modeling can all be framed as generative steps in a single cosmological forward-modeling chain. This suggests that the term has evolved from naming a specific framework toward denoting a methodological style centered on conditional generation, calibration to simulations or data, and scalable mock production.

## 6. Distinct CMS usage: the cosmic muon generator

Outside cosmology, “CosmoGen” refers in CMS to CMSCGEN, the dedicated cosmic muon Monte Carlo event generator for the CMS experiment [1107.0427]. Its purpose is entirely different from the cosmological usages: it provides a fast, parametrized description of the cosmic-ray muon flux at the Earth’s surface and propagates muons through the underground environment of CMS, located approximately 90 m below the surface [1107.0427].

The surface flux is parameterized in terms of momentum $p$, zenith angle $\theta$, and azimuth $\phi$. Defining
$$
L=\log_{10}\!\left(\frac{p}{\mathrm{GeV}}\right),
$$
the momentum-shape function is
$$
s(L)=a_0+a_1L+a_2L^2+a_3L^3+a_4L^4+a_5L^5+a_6L^6,
$$
and the angular modulation is
$$
z(\cos\theta,L)=b_0(L)+b_1(L)\cos\theta+b_2(L)\cos^2\theta.
$$
Assuming azimuthal isotropy,
$$
\frac{d\Phi}{dp\,d\cos\theta\,d\phi}
= C_{\mathrm{norm}}\frac{1}{p^3}s(L)\,z(\cos\theta,L)\,\frac{1}{2\pi},
$$
with $C_{\mathrm{norm}}$ chosen to reproduce the measured vertical muon flux at $p=100$ GeV [1107.0427].

CMSCGEN supports both surface generation followed by underground propagation and direct underground generation. It models the site-specific material map, including concrete structures, access shafts, the collision and service caverns, parts of the LHC tunnel, and two average-density geological layers, and computes the water-equivalent slant depth to determine energy loss before handing particles to the full GEANT4-based detector simulation [1107.0427]. Validation used approximately 300 million cosmic events recorded in 2008 with the CMS solenoid at $B=3.8$ T; the simulated spectrum, falling approximately as a power law with index $-2.7$, was reported to describe the measured energy spectrum well, and shaft-correlated intensity enhancements were reproduced qualitatively [1107.0427].

The only substantive connection between CMSCGEN and cosmological CosmoGen frameworks is nominal. They address different physical systems, operate in different software ecosystems, and solve different inference or simulation problems.

## 7. Limitations, interpretation, and future directions

The principal limitations of the cosmological model generator are internal to its symbolic-regression design. Tree-based genetic programming redundantly explores semantically equivalent expressions; short generation-stage chains can overfit the targeted objective; and the operator grammar constrains expressivity [2509.15453]. The framework therefore does not eliminate the need for downstream posterior exploration, evidence calculation, or theory-motivated physical interpretation. Its own authors explicitly identify exhaustive symbolic regression, expanded primitives, broader datasets, and more principled complexity penalties such as Minimum Description Length as natural next steps [2509.15453].

The broader pipeline interpretation inherits additional module-specific constraints. genetIC’s filters must balance ringing against aliasing, and extremely strong constraints yield high $\Delta\chi^2$ realizations [2006.01841]. CosmoMIA requires calibration and is demonstrated at $z=1.1$ under a specified cosmology and plane-parallel RSD treatment [2402.17581]. CosmiXs is limited to 5 GeV–100 TeV, standard two-body channels, and shower-model assumptions that exclude model-dependent ISR from dark matter [2312.01153]. CosmoGLINT can suffer from domain shift outside its training distribution and currently models per-step properties with independent softmax heads rather than an explicit joint density [2506.16843].

These limitations do not undermine the underlying conceptual shift. The collective literature indicates a movement from manually specified cosmological ansätze and purely handcrafted mock-making recipes toward data-informed generative systems that combine symbolic search, approximate dynamics, conditional sampling, and modular validation. In the narrow sense of [2509.15453], CosmoGen is a framework for machine-assisted cosmological model discovery. In the broader sense emerging across recent work, it names a generative paradigm for cosmological computation.

Source: https://www.emergentmind.com/topics/cosmogen