Papers
Topics
Authors
Recent
Search
2000 character limit reached

Flagship & GAEA Galaxy Mock Catalogues

Updated 22 January 2026
  • Flagship and GAEA galaxy mock catalogues are advanced synthetic datasets that combine gravity-only N-body simulations with semi-analytic models to reproduce realistic galaxy populations.
  • They facilitate precision cosmology by enabling end-to-end survey validation, forward modelling of observables, and testing of cosmological models across cosmic time.
  • These catalogues integrate detailed methodologies—from light-cone construction to baryonic process modelling—to support multiwavelength surveys such as Euclid and CSST.

Flagship and GAEA galaxy mock catalogues are major computational products designed to facilitate the interpretation, calibration, and exploitation of next-generation extragalactic surveys including Euclid and the China Space Station Telescope (CSST). The Flagship catalogue is a phenomenological, gravity-only mock built primarily for the Euclid mission with the aim of supporting precision cosmological and weak-lensing analyses on the survey scale. The GAEA-based catalogues are physically motivated, semi-analytic mocks constructed atop detailed merger trees from large N-body runs and tuned to model the baryonic physics central to galaxy formation. Together, these catalogues define the state of the art for synthetic galaxy populations in cosmological contexts, enabling rigorous end-to-end validation of survey pipelines, forward-modelling of observables, and testing of theoretical models across cosmic time.

1. Simulation Frameworks: N-body and Merger Trees

The Euclid Flagship mock catalogue is generated from the “Flagship 2” (FS2) N-body run, with 16,00034×101216{,}000^3 \approx 4 \times 10^{12} dark matter particles within a 3600h13600\,h^{-1} Mpc periodic box, achieving a particle mass of mp=1×109h1Mm_p = 1 \times 10^9\, h^{-1}\, M_\odot and gravitational softening of 4.5h14.5\, h^{-1} kpc. Cosmological parameters are set to Ωm=0.319\Omega_m = 0.319, Ωb=0.049\Omega_b = 0.049, h=0.67h = 0.67, ns=0.96n_s = 0.96, and σ80.813\sigma_8 \approx 0.813. A full-sky lightcone is produced “on the fly” by recording particle crossings of the observer's past light cone, yielding \sim31 trillion records and 3600h13600\,h^{-1}0700 TB of data; the analysed octant contains 16 billion haloes up to 3600h13600\,h^{-1}1 (Collaboration et al., 2024).

GAEA-based catalogues, such as those for CSST, are constructed atop the Jiutian N-body runs. Two simulations are employed: Jiutian-1G (3600h13600\,h^{-1}2, 3600h13600\,h^{-1}3) and Jiutian-2G (3600h13600\,h^{-1}4, 3600h13600\,h^{-1}5), each with 3600h13600\,h^{-1}6 particles and Planck2018 cosmology (3600h13600\,h^{-1}7, 3600h13600\,h^{-1}8, 3600h13600\,h^{-1}9). Halos are identified with a friends-of-friends (FOF) algorithm, and subhalos/merger trees are extracted using the HBT+ algorithm, which robustly tracks self-bound remnants across cosmic time and mitigates “overmerging” found in position-space-only approaches (Tan et al., 5 Nov 2025).

2. Galaxy Population Assignment: HOD/Abundance Matching versus Semi-Analytic Models

The Flagship catalogue populates halos with galaxies through a combination of Halo Occupation Distribution (HOD) modelling and abundance matching. Centrals are assigned according to mp=1×109h1Mm_p = 1 \times 10^9\, h^{-1}\, M_\odot0 for mp=1×109h1Mm_p = 1 \times 10^9\, h^{-1}\, M_\odot1; satellites follow mp=1×109h1Mm_p = 1 \times 10^9\, h^{-1}\, M_\odot2 with mp=1×109h1Mm_p = 1 \times 10^9\, h^{-1}\, M_\odot3 and mp=1×109h1Mm_p = 1 \times 10^9\, h^{-1}\, M_\odot4. The cumulative galaxy function is

mp=1×109h1Mm_p = 1 \times 10^9\, h^{-1}\, M_\odot5

Luminosities are assigned via de-scattered cumulative luminosity functions (GOODS/SDSS mp=1×109h1Mm_p = 1 \times 10^9\, h^{-1}\, M_\odot6-band) and double power-law or Schechter-like conditional luminosity functions, with explicit scatter for centrals (mp=1×109h1Mm_p = 1 \times 10^9\, h^{-1}\, M_\odot7). Satellite luminosities are determined with a halo-dependent CLF. The approach is calibrated to reproduce low-mp=1×109h1Mm_p = 1 \times 10^9\, h^{-1}\, M_\odot8 luminosity functions and two-point clustering, ensuring consistency with observed number densities and spatial correlations (Collaboration et al., 2024).

In contrast, the GAEA-based mock employs the GAEA semi-analytic model to predict the internal and observable properties of galaxies from first principles. Physical processes encoded include radiative gas cooling, an mp=1×109h1Mm_p = 1 \times 10^9\, h^{-1}\, M_\odot9-based star formation law (4.5h14.5\, h^{-1}0), chemical enrichment with delayed recycling, redshift-dependent stellar feedback, AGN radio/cold accretion, and satellite stripping. The relevant evolution equations include

4.5h14.5\, h^{-1}1

The differential equations are parameterized and calibrated against local stellar mass functions (Li & White 2009), 4.5h14.5\, h^{-1}2–4.5h14.5\, h^{-1}3 scaling, and HI/quenched fractions, with explicit treatment of orphan galaxies after subhalo disruption and merging timescale estimation via dynamical friction (Tan et al., 5 Nov 2025).

3. Baryonic and Photometric Properties: SEDs, Dust, Emission Lines

Flagship assigns galaxy photometric and structural observables post hoc using template matching and empirical scaling relations. For each mock galaxy, apparent magnitudes are computed in 30 bands (Euclid VIS/NISP, plus broad SED interpolation among 136 COSMOS templates and extinction by Prevot/Calzetti laws). Structural properties include Sérsic bulge and exponential disk parameters (e.g., 4.5h14.5\, h^{-1}4), concentration, bulge fraction, triaxial axis ratios, color–4.5h14.5\, h^{-1}5 relationships, and SFR from UV-based conversions (Kennicutt 1998). Emission lines—H4.5h14.5\, h^{-1}6, H4.5h14.5\, h^{-1}7, [OII], [OIII], [NII], [SII], [SIII]—are assigned using SFR and metallicity-based prescriptions, including BPT diagram evolution and scatter. Lensing parameters (convergence, shear, deflection) are assigned from all-sky HEALPix “onion” maps at 4.5h14.5\, h^{-1}8 (Collaboration et al., 2024).

The GAEA catalogue computes SEDs and magnitudes using StarDuster, a neural-network-based module trained on radiative transfer (SKIRT) simulations to model dust attenuation as a function of galaxy geometry, mass, and inclination. Key variables include dust mass 4.5h14.5\, h^{-1}9 and dust optical depths Ωm=0.319\Omega_m = 0.3190, Ωm=0.319\Omega_m = 0.3191, with “birth-cloud” attenuation for stellar populations Ωm=0.319\Omega_m = 0.3192 Myr old. The attenuated luminosity is expressed as

Ωm=0.319\Omega_m = 0.3193

The model supports full broadband SED prediction (UV through far-IR), including spatially resolved dust geometry effects. The GAEA mocks do not provide emission line predictions by default; these require further post-processing (Tan et al., 5 Nov 2025).

4. Light-Cone Construction and Survey Selection

In Flagship, galaxies are placed within the observer’s past lightcone over one octant to Ωm=0.319\Omega_m = 0.3194 using dark matter particle outputs as they cross the lightcone surface. Halo identification is performed with ROCKSTAR on overlapping “bricks,” followed by galaxy assignment and full-sky projection. The magnitude-limited sample, Ωm=0.319\Omega_m = 0.3195, includes 3.4 billion galaxies, with completeness to Ωm=0.319\Omega_m = 0.3196 (haloes down to 10 particle limit), exceeding Euclid Wide Survey depth (Ωm=0.319\Omega_m = 0.3197, 10Ωm=0.319\Omega_m = 0.3198 point-source) (Collaboration et al., 2024).

The GAEA-based lightcone is constructed via the Blic code, which aligns simulation volumes along the line of sight and interpolates positions, velocities, stellar masses, and magnitudes between 128 output snapshots using cubic or linear schemes. The catalogue supports CSST survey footprints: deep (Ωm=0.319\Omega_m = 0.3199 degΩb=0.049\Omega_b = 0.0490), ultra-deep (Ωb=0.049\Omega_b = 0.0491 degΩb=0.049\Omega_b = 0.0492), and variable band selection. Catalog completeness in the deep survey (seven bands) reaches 90% cumulative at Ωb=0.049\Omega_b = 0.0493 (extended to Ωb=0.049\Omega_b = 0.0494 in ultra-deep); selection in Ωb=0.049\Omega_b = 0.0495 alone raises Ωb=0.049\Omega_b = 0.0496 to Ωb=0.049\Omega_b = 0.04972.1 (Tan et al., 5 Nov 2025).

5. Validation, Observational Agreement, and Cosmological Consistency

Validation in the Flagship mock is extensive, encompassing weak lensing, clustering, and internal galaxy statistics. The convergence power spectrum Ωb=0.049\Omega_b = 0.0498 matches Halofit and the Euclid Emulator2 to better than 5% for Ωb=0.049\Omega_b = 0.0499; shear 2-point and 3-point statistics agree with analytic predictions to 10%. Redshift-space multipoles h=0.67h = 0.670 and real-space 2PCFs are reproduced within 1h=0.67h = 0.671, enabling robust extraction of linear bias and Alcock–Paczynski parameters. Galaxy occupation and distribution in clusters reproduces HOD expectations and NFW profile stacking, while cluster luminosity functions and color–magnitude diagrams recover observed bimodality and faint-end slopes. Comparisons with the GAEA model (e.g., cluster velocity dispersions) show agreement at the h=0.67h = 0.67210% level for relevant observables (Collaboration et al., 2024).

GAEA mocks are validated against stellar mass functions (Li & White 2009, GAMA) at h=0.67h = 0.673 and h=0.67h = 0.674, luminosity functions (SDSS ugriz), gas fractions (h=0.67h = 0.675), half-mass size–mass relations (Shen et al. 2003), and projected 2PCFs versus SDSS. These are reproduced within calibration and systematic uncertainties, except for moderate underprediction at high stellar mass and h=0.67h = 0.676 (attributable to AGN feedback tuning). The mock’s photometric redshift distributions, SED dimming (up to h=0.67h = 0.6771 mag in h=0.67h = 0.678-band from dust), and clustering amplitude are demonstrated to be convergent across simulation resolutions (Tan et al., 5 Nov 2025).

6. Data Accessibility, Applications, and Comparative Properties

The Euclid Flagship catalogue is publicly hosted on CosmoHub in Apache Parquet format, with h=0.67h = 0.6795.9 TB in total and ns=0.96n_s = 0.960 TB for the magnitude-limited galaxy sample. An SQL/Hive metadata interface, Python/ROOT/Parquet readers, and sample Jupyter notebooks are provided. The dataset supports precomputed covariance estimation (100 internal jackknife patches) and is integrable with the Euclid Science Ground Segment pipeline for photometric redshift, shape measurement, and slitless spectroscopy validation. Wider applications include cosmological forecasts for DESI, LSST, Roman, and end-to-end survey systematics studies (Collaboration et al., 2024).

The GAEA-based output is oriented toward forward-modeling galaxy evolution, calibration of CSST photometric and SFR indicators, and multiwavelength studies. It is particularly suited to studies requiring physically consistent SEDs, dust, gas, and sizes, enabling joint analyses with Flagship for lensing/clustering covariance. However, emission line predictions and observational systematics (e.g., blending, detection incompleteness) must be appended in downstream processing. Future extensions include high-ns=0.96n_s = 0.961 AGN feedback improvements, line emission modeling, and photometric error pipelines (Tan et al., 5 Nov 2025).

A summary of key differences appears in the table below:

Catalogue Simulation Volume/Res. Galaxy Assignment Main Strengths
Flagship (FS2) ns=0.96n_s = 0.962 Mpc, ns=0.96n_s = 0.963 HOD + abundance matching Lensing, clustering, completeness, lightcone fidelity
GAEA (Jiutian) ns=0.96n_s = 0.964–ns=0.96n_s = 0.965, ns=0.96n_s = 0.966–ns=0.96n_s = 0.967 Semi-analytic (baryonic) SED realism, physical gas/stars, merger history

7. Scientific and Methodological Implications

The Flagship and GAEA catalogues represent complementary paradigms in large-scale galaxy mock construction. The Flagship approach—phenomenological, optimized for cosmological signal extraction—yields unparalleled volume, statistical power, and self-consistent lensing properties. The GAEA semi-analytic machinery enables explicit connection between observable quantities and underlying baryonic processes, at some cost in direct cosmological volume and resolution.

This contrast enables targeted usage: Flagship for precision cosmology, survey systematics, and end-to-end mock pipelines; GAEA for galaxy evolution, quenching pathways, scaling relation studies, and multiwavelength observables. Joint analyses, in which GAEA’s physical realism is used to interpret or reweight Flagship mocks, represent a promising avenue for fully leveraging multi-survey data. Future work will likely merge these approaches further through the inclusion of improved baryonic physics in large-volume, lightcone-enabled N-body backgrounds, and more sophisticated treatment of hydrodynamical and AGN feedback effects (Collaboration et al., 2024, Tan et al., 5 Nov 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Flagship and GAEA Galaxy Mock Catalogues.