Papers
Topics
Authors
Recent
Search
2000 character limit reached

CosmoGen: Cosmological Model Generation

Updated 12 July 2026
  • CosmoGen is a computational framework that uses evolutionary symbolic regression and Bayesian inference to discover interpretable dark-energy models addressing H₀ and S₈ tensions.
  • It integrates a genetic programming engine with cosmological solvers like CLASS and MontePython to evaluate candidate models based on chi-squared minimization and complexity penalties.
  • The broader ecosystem includes modular pipelines for initial conditions, halo catalog synthesis, source spectra, and even cosmic muon generation in CMS, highlighting its versatile applications.

CosmoGen most specifically denotes a cosmological model generator that couples evolutionary symbolic regression to Boltzmann and Bayesian inference codes in order to generate interpretable dark-energy models targeted at specific observational goals, notably the alleviation of the H0H_0 and S8S_8 tensions (Castelão et al., 18 Sep 2025). In adjacent literature, the same label also appears as an umbrella for cosmological generative pipelines spanning initial-condition construction, mock-halo synthesis, indirect-detection source spectra, and line-intensity-mapping catalog generation, and in CMS it is used colloquially for the cosmic-muon event generator CMSCGEN (Stopyra et al., 2020, Forero-Sánchez et al., 2024, Arina et al., 2023, Moriwaki et al., 20 Jun 2025, Sonnenschein, 2011).

1. Nomenclature and scope

The term “CosmoGen” is not univocal. In its strictest sense, it names the framework introduced in “CosmoGen: a cosmological model generator” (Castelão et al., 18 Sep 2025), where the objective is automated model discovery in precision cosmology. There, candidate dark-energy fluids are generated symbolically, inserted into a modified cosmological solver, and ranked against data.

A second usage is infrastructural rather than nominal. Several recent works describe modules that either serve a “CosmoGen pipeline” or are presented explicitly in the context of cosmological generative modeling: genetIC for constrained and genetically modified zoom initial conditions (Stopyra et al., 2020), CosmoMIA for converting coarse-grid halo counts into tracer coordinates and velocities in real and redshift space (Forero-Sánchez et al., 2024), CosmiXs for tabulated dark-matter source spectra used in indirect-detection source terms (Arina et al., 2023), and CosmoGLINT for Transformer-based generation of galaxy populations from dark-matter-only halo catalogs (Moriwaki et al., 20 Jun 2025). This suggests a broader usage in which “CosmoGen” denotes an end-to-end computational ecosystem for cosmological forward modeling rather than a single executable.

A third usage is domain-specific and unrelated to cosmological model discovery. Within CMS, “CosmoGen” is a colloquial name for CMSCGEN, the dedicated cosmic muon Monte Carlo event generator used in the CMS detector simulation chain (Sonnenschein, 2011). The coexistence of these meanings is a recurring source of ambiguity.

2. Core framework: evolutionary generation of dark-energy models

In the 2025 usage, CosmoGen is a model-building framework that leverages evolutionary symbolic regression tightly coupled to cosmological computations to automatically generate interpretable dark-energy models (Castelão et al., 18 Sep 2025). Its proof-of-concept problem is narrowly defined: generate homogeneous dark-energy fluid models that can alleviate the cosmological tensions in H0H_0 and S8S_8 without abandoning a direct confrontation with CMB and weak-lensing information.

The generated models are represented through a normalized dark-energy density,

ΩDE(a)=Af(a;D),\Omega_{\mathrm{DE}}(a) = \frac{A}{f(a;D)},

where aa is the scale factor, DD is a free parameter, and AA is fixed by closure. Using

ΩDE,0=1−∑iΩi−Ωk,\Omega_{\mathrm{DE},0} = 1 - \sum_i \Omega_i - \Omega_k,

the framework sets

A=f(1;D) ΩDE,0,A = f(1;D)\,\Omega_{\mathrm{DE},0},

so that

S8S_80

The symbolic forms for S8S_81 are produced by genetic programming using the grammar with operators S8S_82 acting on terminals S8S_83 (Castelão et al., 18 Sep 2025).

This representation is deliberately constrained. It keeps the search space interpretable, preserves direct control over the background expansion history, and permits immediate insertion into a modified CLASS run. A plausible implication is that the framework is designed less as an unconstrained symbolic-search engine than as a targeted mechanism for constructing viable analytic deformations of S8S_84CDM.

3. Evolutionary loop, viability filters, and scoring

CosmoGen integrates a genetic-programming engine with CLASS and MontePython so that each candidate symbolic expression is turned into a working cosmological component, tested numerically, and scored against a tension-aware likelihood (Castelão et al., 18 Sep 2025). The initialization uses a population size of 4096 random expressions, with a maximum initial complexity of 7 nodes. The evolutionary loop runs for 8 generations, with S8S_85 parents, S8S_86 children, mate and mutate probabilities of S8S_87, and a maximum complexity increase per generation of 2.

Before full evaluation, candidates undergo simplification and pre-testing. SymPy and custom routines are used to simplify expressions, transform them into an energy density, test uniqueness, and reject numerical pathologies such as divergences and complex values over representative S8S_88 values S8S_89. Physical viability heuristics require H0H_00 in the broad range H0H_01 and small adiabatic sound speed H0H_02; the best trial values of H0H_03 are stored to seed the MCMC (Castelão et al., 18 Sep 2025).

The candidate is then inserted on-the-fly into a modified CLASS in which H0H_04 is replaced by a homogeneous dark-energy fluid with H0H_05. MontePython performs short Metropolis–Hastings chains in the reduced parameter space H0H_06, with all other parameters fixed to Planck 2018 H0H_07CDM means during the generation stage. The likelihood combines Planck 2018 high-H0H_08 TTTEEE (lite), low-H0H_09 EE, and low-S8S_80 TT likelihoods with Gaussian priors on SH0ES S8S_81 and KiDS-1000 S8S_82 (Castelão et al., 18 Sep 2025).

The optimization target is not the raw minimum S8S_83 alone. CosmoGen defines

S8S_84

where S8S_85 is the expression-tree complexity. Since typical S8S_86 values are S8S_87–S8S_88 and S8S_89 is ΩDE(a)=Af(a;D),\Omega_{\mathrm{DE}}(a) = \frac{A}{f(a;D)},0–ΩDE(a)=Af(a;D),\Omega_{\mathrm{DE}}(a) = \frac{A}{f(a;D)},1, the complexity penalty acts as a tiebreaker rather than as the dominant regularizer (Castelão et al., 18 Sep 2025). Tournament selection then chooses parents for the next generation.

Component Setting Value
Initial population Random expressions 4096
Generations Evolution limit 8
Parent/child strategy ΩDE(a)=Af(a;D),\Omega_{\mathrm{DE}}(a) = \frac{A}{f(a;D)},2 ΩDE(a)=Af(a;D),\Omega_{\mathrm{DE}}(a) = \frac{A}{f(a;D)},3, ΩDE(a)=Af(a;D),\Omega_{\mathrm{DE}}(a) = \frac{A}{f(a;D)},4
Mate/mutate probabilities Per offspring step 0.5 / 0.5
Initial complexity cap Generation 1 7 nodes
Complexity growth cap Per generation 2 nodes

The associated cosmological computations are standard but central to the loop. For the implemented flat case,

ΩDE(a)=Af(a;D),\Omega_{\mathrm{DE}}(a) = \frac{A}{f(a;D)},5

From the continuity equation,

ΩDE(a)=Af(a;D),\Omega_{\mathrm{DE}}(a) = \frac{A}{f(a;D)},6

CosmoGen derives

ΩDE(a)=Af(a;D),\Omega_{\mathrm{DE}}(a) = \frac{A}{f(a;D)},7

CLASS then computes ΩDE(a)=Af(a;D),\Omega_{\mathrm{DE}}(a) = \frac{A}{f(a;D)},8, ΩDE(a)=Af(a;D),\Omega_{\mathrm{DE}}(a) = \frac{A}{f(a;D)},9, aa0, and derived parameters including

aa1

This makes the search intrinsically physics-aware rather than purely symbolic (Castelão et al., 18 Sep 2025).

4. Generated models, the CG example, and empirical results

The pipeline yields aa2 viable final models, and the highest-scoring expressions are compact. The illustrative forms reported include

aa3

(Castelão et al., 18 Sep 2025). Many of the top candidates therefore combine logarithms with low-degree polynomials in aa4 and aa5, consistent with the framework’s built-in preference for interpretable expressions.

The representative “CG model” selected for deeper analysis has

aa6

with equation of state

aa7

The reported interpretation is that for aa8, aa9 can cross the phantom divide and return, whereas for DD0, the evolution is slower and less extreme (Castelão et al., 18 Sep 2025).

The final Bayesian assessment differs from the tension-aware generation stage. It uses standard datasets without explicit DD1 or DD2 priors: Planck 2018 and KiDS–VIKING KV-450. The varied parameters are DD3, DD4, DD5, DD6, DD7, and DD8, with derived DD9 or AA0, plus nuisance parameters AA1 and AA2. Spatial flatness is assumed (Castelão et al., 18 Sep 2025).

Dataset CG model AA3CDM
Planck 2018 AA4 AA5 AA6
Planck 2018 AA7 AA8 AA9
KV-450 ΩDE,0=1−∑iΩi−Ωk,\Omega_{\mathrm{DE},0} = 1 - \sum_i \Omega_i - \Omega_k,0 ΩDE,0=1−∑iΩi−Ωk,\Omega_{\mathrm{DE},0} = 1 - \sum_i \Omega_i - \Omega_k,1 ΩDE,0=1−∑iΩi−Ωk,\Omega_{\mathrm{DE},0} = 1 - \sum_i \Omega_i - \Omega_k,2
KV-450 ΩDE,0=1−∑iΩi−Ωk,\Omega_{\mathrm{DE},0} = 1 - \sum_i \Omega_i - \Omega_k,3 ΩDE,0=1−∑iΩi−Ωk,\Omega_{\mathrm{DE},0} = 1 - \sum_i \Omega_i - \Omega_k,4 ΩDE,0=1−∑iΩi−Ωk,\Omega_{\mathrm{DE},0} = 1 - \sum_i \Omega_i - \Omega_k,5
Planck–KV-450 ΩDE,0=1−∑iΩi−Ωk,\Omega_{\mathrm{DE},0} = 1 - \sum_i \Omega_i - \Omega_k,6 tension ΩDE,0=1−∑iΩi−Ωk,\Omega_{\mathrm{DE},0} = 1 - \sum_i \Omega_i - \Omega_k,7 ΩDE,0=1−∑iΩi−Ωk,\Omega_{\mathrm{DE},0} = 1 - \sum_i \Omega_i - \Omega_k,8

The reported evidence remains mildly unfavorable to the generated model: ΩDE,0=1−∑iΩi−Ωk,\Omega_{\mathrm{DE},0} = 1 - \sum_i \Omega_i - \Omega_k,9 for Planck 2018 and A=f(1;D) ΩDE,0,A = f(1;D)\,\Omega_{\mathrm{DE},0},0 for KV-450, both indicating weak preference for A=f(1;D) ΩDE,0,A = f(1;D)\,\Omega_{\mathrm{DE},0},1CDM (Castelão et al., 18 Sep 2025). Nonetheless, the framework demonstrates that a machine-generated analytic fluid can shift the Planck-inferred A=f(1;D) ΩDE,0,A = f(1;D)\,\Omega_{\mathrm{DE},0},2 upward by approximately A=f(1;D) ΩDE,0,A = f(1;D)\,\Omega_{\mathrm{DE},0},3 and reduce the inter-dataset A=f(1;D) ΩDE,0,A = f(1;D)\,\Omega_{\mathrm{DE},0},4 tension from A=f(1;D) ΩDE,0,A = f(1;D)\,\Omega_{\mathrm{DE},0},5 to A=f(1;D) ΩDE,0,A = f(1;D)\,\Omega_{\mathrm{DE},0},6.

A common misconception is that CosmoGen is presented as a wholesale replacement for model comparison. The reported results do not support that interpretation. The Bayesian evidence still weakly prefers A=f(1;D) ΩDE,0,A = f(1;D)\,\Omega_{\mathrm{DE},0},7CDM, and the generation-stage likelihood uses targeted priors explicitly intended to encourage tension alleviation (Castelão et al., 18 Sep 2025). A more precise characterization is that CosmoGen is a model-discovery engine for proposing interpretable extensions that merit subsequent conventional inference.

5. CosmoGen as a broader computational ecosystem

In several neighboring contexts, “CosmoGen” functions as an umbrella for modular cosmological generation workflows rather than for symbolic model discovery alone. The modules described in that literature span initial conditions, halo catalog generation, source spectra, and galaxy-population synthesis.

genetIC provides cosmological initial conditions for A=f(1;D) ΩDE,0,A = f(1;D)\,\Omega_{\mathrm{DE},0},8-body simulations, including nested zoom regions and “genetic modifications” that alter halo history, mass, or environment while maintaining consistency with the statistical ensemble (Stopyra et al., 2020). It constructs Gaussian random fields using a white-noise representation,

A=f(1;D) ΩDE,0,A = f(1;D)\,\Omega_{\mathrm{DE},0},9

and computes 1LPT displacements

S8S_800

For multi-resolution zooms it introduces complementary Fourier filters,

S8S_801

with default choices S8S_802 and S8S_803 (Stopyra et al., 2020). The code reports sub-percent precision for modifications within zoom initial conditions and avoids the large ghost-padding memory cost of traditional nested-grid convolution schemes.

CosmoMIA addresses a different stage of the forward model: it takes coarse-grid dark-matter realizations and target number counts per cell and generates coordinates and velocities for individual tracers (Forero-Sánchez et al., 2024). Its four stages are tracer assignment, attractor identification, tracer collapse, and redshift-space distortions. At S8S_804, it reproduces AbacusSummit halo clustering with the power-spectrum monopole within S8S_805 up to S8S_806 in real space, within S8S_807 to S8S_808 and S8S_809 to S8S_810 in redshift space, and the quadrupole within S8S_811 to S8S_812 (Forero-Sánchez et al., 2024). The coherent velocity field is enhanced through

S8S_813

with fiducial S8S_814 and S8S_815, before mapping to redshift space via

S8S_816

(Forero-Sánchez et al., 2024).

CosmiXs supplies tabulated source spectra for indirect dark-matter searches and is explicitly described as ready to plug into “CosmoGen” pipelines (Arina et al., 2023). It tabulates S8S_817 for S8S_818, S8S_819, S8S_820, and S8S_821 for 64 dark-matter masses between 5 GeV and 100 TeV and for standard fermionic and bosonic two-body final states, including helicity- and polarization-resolved variants (Arina et al., 2023). The conversion to injection spectra is

S8S_822

For annihilation and decay, the source terms are

S8S_823

respectively (Arina et al., 2023).

CosmoGLINT extends the generative interpretation of “CosmoGen” to galaxy populations and line-intensity-mapping mocks (Moriwaki et al., 20 Jun 2025). It is a decoder-only Transformer with 4 layers, 8 heads, hidden size 128, and 100-bin softmax heads per property, trained on IllustrisTNG to generate variable-length galaxy sequences conditioned on halo mass. For each halo it emits tuples

S8S_824

with factorized conditional probability

S8S_825

under causal masking (Moriwaki et al., 20 Jun 2025). At S8S_826, training uses halos with S8S_827 and galaxies with S8S_828, yielding 330,287 halos and 5,027,636 galaxies; the reported generation throughput is about 250,000 galaxies/sec or about 16,000 halos/sec (Moriwaki et al., 20 Jun 2025). It reproduces the voxel intensity distribution and the power spectrum in both real and redshift space and is trained across 15 snapshots from S8S_829 to S8S_830 for lightcone generation (Moriwaki et al., 20 Jun 2025).

Taken together, these modules define a broad computational interpretation of CosmoGen: initial conditions, approximate gravity, tracer realization, galaxy population synthesis, and particle-physics source modeling can all be framed as generative steps in a single cosmological forward-modeling chain. This suggests that the term has evolved from naming a specific framework toward denoting a methodological style centered on conditional generation, calibration to simulations or data, and scalable mock production.

6. Distinct CMS usage: the cosmic muon generator

Outside cosmology, “CosmoGen” refers in CMS to CMSCGEN, the dedicated cosmic muon Monte Carlo event generator for the CMS experiment (Sonnenschein, 2011). Its purpose is entirely different from the cosmological usages: it provides a fast, parametrized description of the cosmic-ray muon flux at the Earth’s surface and propagates muons through the underground environment of CMS, located approximately 90 m below the surface (Sonnenschein, 2011).

The surface flux is parameterized in terms of momentum S8S_831, zenith angle S8S_832, and azimuth S8S_833. Defining

S8S_834

the momentum-shape function is

S8S_835

and the angular modulation is

S8S_836

Assuming azimuthal isotropy,

S8S_837

with S8S_838 chosen to reproduce the measured vertical muon flux at S8S_839 GeV (Sonnenschein, 2011).

CMSCGEN supports both surface generation followed by underground propagation and direct underground generation. It models the site-specific material map, including concrete structures, access shafts, the collision and service caverns, parts of the LHC tunnel, and two average-density geological layers, and computes the water-equivalent slant depth to determine energy loss before handing particles to the full GEANT4-based detector simulation (Sonnenschein, 2011). Validation used approximately 300 million cosmic events recorded in 2008 with the CMS solenoid at S8S_840 T; the simulated spectrum, falling approximately as a power law with index S8S_841, was reported to describe the measured energy spectrum well, and shaft-correlated intensity enhancements were reproduced qualitatively (Sonnenschein, 2011).

The only substantive connection between CMSCGEN and cosmological CosmoGen frameworks is nominal. They address different physical systems, operate in different software ecosystems, and solve different inference or simulation problems.

7. Limitations, interpretation, and future directions

The principal limitations of the cosmological model generator are internal to its symbolic-regression design. Tree-based genetic programming redundantly explores semantically equivalent expressions; short generation-stage chains can overfit the targeted objective; and the operator grammar constrains expressivity (Castelão et al., 18 Sep 2025). The framework therefore does not eliminate the need for downstream posterior exploration, evidence calculation, or theory-motivated physical interpretation. Its own authors explicitly identify exhaustive symbolic regression, expanded primitives, broader datasets, and more principled complexity penalties such as Minimum Description Length as natural next steps (Castelão et al., 18 Sep 2025).

The broader pipeline interpretation inherits additional module-specific constraints. genetIC’s filters must balance ringing against aliasing, and extremely strong constraints yield high S8S_842 realizations (Stopyra et al., 2020). CosmoMIA requires calibration and is demonstrated at S8S_843 under a specified cosmology and plane-parallel RSD treatment (Forero-Sánchez et al., 2024). CosmiXs is limited to 5 GeV–100 TeV, standard two-body channels, and shower-model assumptions that exclude model-dependent ISR from dark matter (Arina et al., 2023). CosmoGLINT can suffer from domain shift outside its training distribution and currently models per-step properties with independent softmax heads rather than an explicit joint density (Moriwaki et al., 20 Jun 2025).

These limitations do not undermine the underlying conceptual shift. The collective literature indicates a movement from manually specified cosmological ansätze and purely handcrafted mock-making recipes toward data-informed generative systems that combine symbolic search, approximate dynamics, conditional sampling, and modular validation. In the narrow sense of (Castelão et al., 18 Sep 2025), CosmoGen is a framework for machine-assisted cosmological model discovery. In the broader sense emerging across recent work, it names a generative paradigm for cosmological computation.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to CosmoGen.