CosmoGen: Cosmological Model Generation
- CosmoGen is a computational framework that uses evolutionary symbolic regression and Bayesian inference to discover interpretable dark-energy models addressing H₀ and S₈ tensions.
- It integrates a genetic programming engine with cosmological solvers like CLASS and MontePython to evaluate candidate models based on chi-squared minimization and complexity penalties.
- The broader ecosystem includes modular pipelines for initial conditions, halo catalog synthesis, source spectra, and even cosmic muon generation in CMS, highlighting its versatile applications.
CosmoGen most specifically denotes a cosmological model generator that couples evolutionary symbolic regression to Boltzmann and Bayesian inference codes in order to generate interpretable dark-energy models targeted at specific observational goals, notably the alleviation of the and tensions (Castelão et al., 18 Sep 2025). In adjacent literature, the same label also appears as an umbrella for cosmological generative pipelines spanning initial-condition construction, mock-halo synthesis, indirect-detection source spectra, and line-intensity-mapping catalog generation, and in CMS it is used colloquially for the cosmic-muon event generator CMSCGEN (Stopyra et al., 2020, Forero-Sánchez et al., 2024, Arina et al., 2023, Moriwaki et al., 20 Jun 2025, Sonnenschein, 2011).
1. Nomenclature and scope
The term “CosmoGen” is not univocal. In its strictest sense, it names the framework introduced in “CosmoGen: a cosmological model generator” (Castelão et al., 18 Sep 2025), where the objective is automated model discovery in precision cosmology. There, candidate dark-energy fluids are generated symbolically, inserted into a modified cosmological solver, and ranked against data.
A second usage is infrastructural rather than nominal. Several recent works describe modules that either serve a “CosmoGen pipeline” or are presented explicitly in the context of cosmological generative modeling: genetIC for constrained and genetically modified zoom initial conditions (Stopyra et al., 2020), CosmoMIA for converting coarse-grid halo counts into tracer coordinates and velocities in real and redshift space (Forero-Sánchez et al., 2024), CosmiXs for tabulated dark-matter source spectra used in indirect-detection source terms (Arina et al., 2023), and CosmoGLINT for Transformer-based generation of galaxy populations from dark-matter-only halo catalogs (Moriwaki et al., 20 Jun 2025). This suggests a broader usage in which “CosmoGen” denotes an end-to-end computational ecosystem for cosmological forward modeling rather than a single executable.
A third usage is domain-specific and unrelated to cosmological model discovery. Within CMS, “CosmoGen” is a colloquial name for CMSCGEN, the dedicated cosmic muon Monte Carlo event generator used in the CMS detector simulation chain (Sonnenschein, 2011). The coexistence of these meanings is a recurring source of ambiguity.
2. Core framework: evolutionary generation of dark-energy models
In the 2025 usage, CosmoGen is a model-building framework that leverages evolutionary symbolic regression tightly coupled to cosmological computations to automatically generate interpretable dark-energy models (Castelão et al., 18 Sep 2025). Its proof-of-concept problem is narrowly defined: generate homogeneous dark-energy fluid models that can alleviate the cosmological tensions in and without abandoning a direct confrontation with CMB and weak-lensing information.
The generated models are represented through a normalized dark-energy density,
where is the scale factor, is a free parameter, and is fixed by closure. Using
the framework sets
so that
0
The symbolic forms for 1 are produced by genetic programming using the grammar with operators 2 acting on terminals 3 (Castelão et al., 18 Sep 2025).
This representation is deliberately constrained. It keeps the search space interpretable, preserves direct control over the background expansion history, and permits immediate insertion into a modified CLASS run. A plausible implication is that the framework is designed less as an unconstrained symbolic-search engine than as a targeted mechanism for constructing viable analytic deformations of 4CDM.
3. Evolutionary loop, viability filters, and scoring
CosmoGen integrates a genetic-programming engine with CLASS and MontePython so that each candidate symbolic expression is turned into a working cosmological component, tested numerically, and scored against a tension-aware likelihood (Castelão et al., 18 Sep 2025). The initialization uses a population size of 4096 random expressions, with a maximum initial complexity of 7 nodes. The evolutionary loop runs for 8 generations, with 5 parents, 6 children, mate and mutate probabilities of 7, and a maximum complexity increase per generation of 2.
Before full evaluation, candidates undergo simplification and pre-testing. SymPy and custom routines are used to simplify expressions, transform them into an energy density, test uniqueness, and reject numerical pathologies such as divergences and complex values over representative 8 values 9. Physical viability heuristics require 0 in the broad range 1 and small adiabatic sound speed 2; the best trial values of 3 are stored to seed the MCMC (Castelão et al., 18 Sep 2025).
The candidate is then inserted on-the-fly into a modified CLASS in which 4 is replaced by a homogeneous dark-energy fluid with 5. MontePython performs short Metropolis–Hastings chains in the reduced parameter space 6, with all other parameters fixed to Planck 2018 7CDM means during the generation stage. The likelihood combines Planck 2018 high-8 TTTEEE (lite), low-9 EE, and low-0 TT likelihoods with Gaussian priors on SH0ES 1 and KiDS-1000 2 (Castelão et al., 18 Sep 2025).
The optimization target is not the raw minimum 3 alone. CosmoGen defines
4
where 5 is the expression-tree complexity. Since typical 6 values are 7–8 and 9 is 0–1, the complexity penalty acts as a tiebreaker rather than as the dominant regularizer (Castelão et al., 18 Sep 2025). Tournament selection then chooses parents for the next generation.
| Component | Setting | Value |
|---|---|---|
| Initial population | Random expressions | 4096 |
| Generations | Evolution limit | 8 |
| Parent/child strategy | 2 | 3, 4 |
| Mate/mutate probabilities | Per offspring step | 0.5 / 0.5 |
| Initial complexity cap | Generation 1 | 7 nodes |
| Complexity growth cap | Per generation | 2 nodes |
The associated cosmological computations are standard but central to the loop. For the implemented flat case,
5
From the continuity equation,
6
CosmoGen derives
7
CLASS then computes 8, 9, 0, and derived parameters including
1
This makes the search intrinsically physics-aware rather than purely symbolic (Castelão et al., 18 Sep 2025).
4. Generated models, the CG example, and empirical results
The pipeline yields 2 viable final models, and the highest-scoring expressions are compact. The illustrative forms reported include
3
(Castelão et al., 18 Sep 2025). Many of the top candidates therefore combine logarithms with low-degree polynomials in 4 and 5, consistent with the framework’s built-in preference for interpretable expressions.
The representative “CG model” selected for deeper analysis has
6
with equation of state
7
The reported interpretation is that for 8, 9 can cross the phantom divide and return, whereas for 0, the evolution is slower and less extreme (Castelão et al., 18 Sep 2025).
The final Bayesian assessment differs from the tension-aware generation stage. It uses standard datasets without explicit 1 or 2 priors: Planck 2018 and KiDS–VIKING KV-450. The varied parameters are 3, 4, 5, 6, 7, and 8, with derived 9 or 0, plus nuisance parameters 1 and 2. Spatial flatness is assumed (Castelão et al., 18 Sep 2025).
| Dataset | CG model | 3CDM |
|---|---|---|
| Planck 2018 4 | 5 | 6 |
| Planck 2018 7 | 8 | 9 |
| KV-450 0 | 1 | 2 |
| KV-450 3 | 4 | 5 |
| Planck–KV-450 6 tension | 7 | 8 |
The reported evidence remains mildly unfavorable to the generated model: 9 for Planck 2018 and 0 for KV-450, both indicating weak preference for 1CDM (Castelão et al., 18 Sep 2025). Nonetheless, the framework demonstrates that a machine-generated analytic fluid can shift the Planck-inferred 2 upward by approximately 3 and reduce the inter-dataset 4 tension from 5 to 6.
A common misconception is that CosmoGen is presented as a wholesale replacement for model comparison. The reported results do not support that interpretation. The Bayesian evidence still weakly prefers 7CDM, and the generation-stage likelihood uses targeted priors explicitly intended to encourage tension alleviation (Castelão et al., 18 Sep 2025). A more precise characterization is that CosmoGen is a model-discovery engine for proposing interpretable extensions that merit subsequent conventional inference.
5. CosmoGen as a broader computational ecosystem
In several neighboring contexts, “CosmoGen” functions as an umbrella for modular cosmological generation workflows rather than for symbolic model discovery alone. The modules described in that literature span initial conditions, halo catalog generation, source spectra, and galaxy-population synthesis.
genetIC provides cosmological initial conditions for 8-body simulations, including nested zoom regions and “genetic modifications” that alter halo history, mass, or environment while maintaining consistency with the statistical ensemble (Stopyra et al., 2020). It constructs Gaussian random fields using a white-noise representation,
9
and computes 1LPT displacements
00
For multi-resolution zooms it introduces complementary Fourier filters,
01
with default choices 02 and 03 (Stopyra et al., 2020). The code reports sub-percent precision for modifications within zoom initial conditions and avoids the large ghost-padding memory cost of traditional nested-grid convolution schemes.
CosmoMIA addresses a different stage of the forward model: it takes coarse-grid dark-matter realizations and target number counts per cell and generates coordinates and velocities for individual tracers (Forero-Sánchez et al., 2024). Its four stages are tracer assignment, attractor identification, tracer collapse, and redshift-space distortions. At 04, it reproduces AbacusSummit halo clustering with the power-spectrum monopole within 05 up to 06 in real space, within 07 to 08 and 09 to 10 in redshift space, and the quadrupole within 11 to 12 (Forero-Sánchez et al., 2024). The coherent velocity field is enhanced through
13
with fiducial 14 and 15, before mapping to redshift space via
16
(Forero-Sánchez et al., 2024).
CosmiXs supplies tabulated source spectra for indirect dark-matter searches and is explicitly described as ready to plug into “CosmoGen” pipelines (Arina et al., 2023). It tabulates 17 for 18, 19, 20, and 21 for 64 dark-matter masses between 5 GeV and 100 TeV and for standard fermionic and bosonic two-body final states, including helicity- and polarization-resolved variants (Arina et al., 2023). The conversion to injection spectra is
22
For annihilation and decay, the source terms are
23
respectively (Arina et al., 2023).
CosmoGLINT extends the generative interpretation of “CosmoGen” to galaxy populations and line-intensity-mapping mocks (Moriwaki et al., 20 Jun 2025). It is a decoder-only Transformer with 4 layers, 8 heads, hidden size 128, and 100-bin softmax heads per property, trained on IllustrisTNG to generate variable-length galaxy sequences conditioned on halo mass. For each halo it emits tuples
24
with factorized conditional probability
25
under causal masking (Moriwaki et al., 20 Jun 2025). At 26, training uses halos with 27 and galaxies with 28, yielding 330,287 halos and 5,027,636 galaxies; the reported generation throughput is about 250,000 galaxies/sec or about 16,000 halos/sec (Moriwaki et al., 20 Jun 2025). It reproduces the voxel intensity distribution and the power spectrum in both real and redshift space and is trained across 15 snapshots from 29 to 30 for lightcone generation (Moriwaki et al., 20 Jun 2025).
Taken together, these modules define a broad computational interpretation of CosmoGen: initial conditions, approximate gravity, tracer realization, galaxy population synthesis, and particle-physics source modeling can all be framed as generative steps in a single cosmological forward-modeling chain. This suggests that the term has evolved from naming a specific framework toward denoting a methodological style centered on conditional generation, calibration to simulations or data, and scalable mock production.
6. Distinct CMS usage: the cosmic muon generator
Outside cosmology, “CosmoGen” refers in CMS to CMSCGEN, the dedicated cosmic muon Monte Carlo event generator for the CMS experiment (Sonnenschein, 2011). Its purpose is entirely different from the cosmological usages: it provides a fast, parametrized description of the cosmic-ray muon flux at the Earth’s surface and propagates muons through the underground environment of CMS, located approximately 90 m below the surface (Sonnenschein, 2011).
The surface flux is parameterized in terms of momentum 31, zenith angle 32, and azimuth 33. Defining
34
the momentum-shape function is
35
and the angular modulation is
36
Assuming azimuthal isotropy,
37
with 38 chosen to reproduce the measured vertical muon flux at 39 GeV (Sonnenschein, 2011).
CMSCGEN supports both surface generation followed by underground propagation and direct underground generation. It models the site-specific material map, including concrete structures, access shafts, the collision and service caverns, parts of the LHC tunnel, and two average-density geological layers, and computes the water-equivalent slant depth to determine energy loss before handing particles to the full GEANT4-based detector simulation (Sonnenschein, 2011). Validation used approximately 300 million cosmic events recorded in 2008 with the CMS solenoid at 40 T; the simulated spectrum, falling approximately as a power law with index 41, was reported to describe the measured energy spectrum well, and shaft-correlated intensity enhancements were reproduced qualitatively (Sonnenschein, 2011).
The only substantive connection between CMSCGEN and cosmological CosmoGen frameworks is nominal. They address different physical systems, operate in different software ecosystems, and solve different inference or simulation problems.
7. Limitations, interpretation, and future directions
The principal limitations of the cosmological model generator are internal to its symbolic-regression design. Tree-based genetic programming redundantly explores semantically equivalent expressions; short generation-stage chains can overfit the targeted objective; and the operator grammar constrains expressivity (Castelão et al., 18 Sep 2025). The framework therefore does not eliminate the need for downstream posterior exploration, evidence calculation, or theory-motivated physical interpretation. Its own authors explicitly identify exhaustive symbolic regression, expanded primitives, broader datasets, and more principled complexity penalties such as Minimum Description Length as natural next steps (Castelão et al., 18 Sep 2025).
The broader pipeline interpretation inherits additional module-specific constraints. genetIC’s filters must balance ringing against aliasing, and extremely strong constraints yield high 42 realizations (Stopyra et al., 2020). CosmoMIA requires calibration and is demonstrated at 43 under a specified cosmology and plane-parallel RSD treatment (Forero-Sánchez et al., 2024). CosmiXs is limited to 5 GeV–100 TeV, standard two-body channels, and shower-model assumptions that exclude model-dependent ISR from dark matter (Arina et al., 2023). CosmoGLINT can suffer from domain shift outside its training distribution and currently models per-step properties with independent softmax heads rather than an explicit joint density (Moriwaki et al., 20 Jun 2025).
These limitations do not undermine the underlying conceptual shift. The collective literature indicates a movement from manually specified cosmological ansätze and purely handcrafted mock-making recipes toward data-informed generative systems that combine symbolic search, approximate dynamics, conditional sampling, and modular validation. In the narrow sense of (Castelão et al., 18 Sep 2025), CosmoGen is a framework for machine-assisted cosmological model discovery. In the broader sense emerging across recent work, it names a generative paradigm for cosmological computation.