---
title: 'EuclidLargeMocks: Galaxy Mocks for Euclid'
url: https://www.emergentmind.com/topics/euclidlargemocks
type: topic
---

# EuclidLargeMocks: Galaxy Mocks for Euclid

EuclidLargeMocks denotes a family of Euclid-like mock spectroscopic catalogues developed for Euclid clustering analyses, with the core Data Release 1-oriented set consisting of 1000 mock galaxy catalogues built from EuclidLargeBox halo light-cones and populated through a halo occupation distribution calibrated on the Euclid Flagship mock galaxy catalogue. Their primary roles are covariance estimation, pipeline validation, robustness tests for two-point statistics, and cosmological parameter inference for the Euclid spectroscopic sample. In the broader Euclid preparation literature, the same mock infrastructure is complemented by Flagship I H\(\alpha\) snapshot catalogues for real-space theory validation, contaminated DR1 catalogues for redshift-interloper studies, relativistic light-cone catalogues for large-scale projection effects, and Flagship 2-based calibrated mock-generation pipelines for wide and deep survey products [2507.12116][2312.00679][2505.04688][2410.00956][2604.14802].

## 1. Catalogue family, scale, and survey targeting

The principal EuclidLargeMocks production is embedded in a two-tier simulation programme. The **Geppetto** set comprises **3500 realizations** from a \(1.2\,h^{-1}\,\mathrm{Gpc}\) box, while the **EuclidLargeBox** set comprises **1000 realizations** from a \(3.38\,h^{-1}\,\mathrm{Gpc}\) box. EuclidLargeMocks are the **1000 galaxy mock realizations** built from EuclidLargeBox and extracted on a **30\(^\circ\)-radius footprint**, corresponding to **2763 deg\(^2\)**. The EuclidLargeBox halo set provides a half-sky output, and the 30\(^\circ\)-radius Euclid footprint is fully contained inside the simulation box, so the derived EuclidLargeMocks are free of replication artifacts. The combined \(3500+1000\) simulation effort is presented as the largest, public set of simulated skies, and the resulting galaxy catalogues are intended for Euclid DR1 galaxy clustering analyses [2507.12116].

| Set | Realizations | Key properties |
|---|---:|---|
| Geppetto | 3500 | \(1.2\,h^{-1}\,\mathrm{Gpc}\) box; 30\(^\circ\) radius; 2763 deg\(^2\); minimum halo mass \(1.52\times10^{11}\,h^{-1}M_\odot\); replication artifacts in Fourier space |
| EuclidLargeBox | 1000 | \(3.38\,h^{-1}\,\mathrm{Gpc}\) box; \(6144^3\) particles; minimum halo mass \(1.48\times10^{11}\,h^{-1}M_\odot\); half-sky halo output; full light-cone volume \(163.86\,h^{-3}\,\mathrm{Gpc}^3\) |
| EuclidLargeMocks | 1000 | Galaxy mocks built from EuclidLargeBox; 30\(^\circ\) radius; 2763 deg\(^2\); comoving volume of a single light-cone \(163.86\,h^{-3}\,\mathrm{Gpc}^3\); total over all realizations \(163{,}862\,h^{-3}\,\mathrm{Gpc}^3\) |

The DR1-specific interloper study uses **1000 synthetic spectroscopic catalogues**, also described as EuclidLargeMocks, whose parent light-cones cover the same **30\(^\circ\)-radius** area and were chosen to be slightly larger than the expected **\(\sim 2500\) deg\(^2\)** DR1 footprint. This continuity of footprint and selection-function design makes the mock family directly usable across covariance, contamination, and inference studies [2505.04688].

## 2. Halo simulation and galaxy-population pipeline

The EuclidLargeMocks halo catalogues are generated with **Pinocchio v5**, a fast approximate halo simulator based on Lagrangian perturbation theory, excursion-set collapse, and fragmentation into haloes. Galaxy catalogues are then built by measuring a **Halo Occupation Distribution (HOD)** from the Euclid Flagship mock, calibrating halo masses between Pinocchio and Flagship, populating haloes with galaxies using the calibrated HOD, and matching the target Euclid spectroscopic selection in \(\mathrm{H}\alpha\) flux [2507.12116].

The HOD is defined in terms of central and satellite occupations measured in bins of halo mass and redshift:
\[
N_{\rm cen}(M_{\rm h}, z\,|\, f_{\ge f_{\rm lim}})=\frac{\#\ \text{of central galaxies}}{\#\ \text{of haloes}},
\]
\[
N_{\rm sat}(M_{\rm h}, z\,|\, f_{\ge f_{\rm lim}})=\frac{\#\ \text{of satellite galaxies}}{\#\ \text{of haloes}}.
\]
Central galaxies are assigned probabilistically using \(N_{\rm cen}\), while satellite counts are Poisson-distributed with mean \(N_{\rm sat}\). Satellites are placed in an NFW profile with velocity model
\[
V_{\rm sat}(r)=f_{\rm v}\,\sqrt{\frac{G M_{\rm h}(<r)}{r}},
\qquad f_{\rm v}=0.7.
\]
To match clustering rather than only abundance, the Pinocchio-to-Flagship halo-mass calibration uses the clustering-matching relation
\[
\log_{10}\!\left(\frac{M_{\rm CM}}{M_\odot}\right)
=
\log_{10}\!\left(\frac{M_{\rm AM}}{M_\odot}\right)
-0.125-0.175(1-z).
\]

For the DR1 interloper catalogues, the intrinsic mock content is augmented with measured-redshift assignment through a calibrated conditional distribution \(P(z_{\rm meas}\,|\,z_{\rm true})\) derived from end-to-end Euclid simulations. Each galaxy carries sky position, true redshift including peculiar velocities, and \(\mathrm{H}\alpha\) flux. The mock selection adopts
\[
f_{\mathrm{H}\alpha} > 10^{-16}\ \mathrm{erg\ s^{-1}\ cm^{-2}},
\]
chosen to emulate the fact that the real Euclid selection is not a sharp flux cut [2505.04688].

A plausible implication is that EuclidLargeMocks are best understood not as a single file-level data product, but as a calibrated mock-production framework with interchangeable observational layers: pure clustering mocks, contaminated redshift catalogues, and specialized light-cone realizations.

## 3. Statistical validation against Flagship and inference equivalence

Validation of EuclidLargeMocks against the Flagship spectroscopic catalogue is performed with **number densities**, **power spectra**, **2-point correlation functions**, and **correlation/covariance matrices**. The number density \(n_g(z)\) is checked for several flux cuts, \(f>0.5f_0\), \(f>f_0\), and \(f>1.5f_0\), and the EuclidLargeMocks reproduce the Flagship number densities well; the reported discrepancies are small and mostly consistent with sample variance. In Fourier space, the monopole, quadrupole, and hexadecapole agree very well on the scales used for standard inference. For
\[
k < 0.2\,h,
\]
Flagship is described as fully consistent with being one realization drawn from EuclidLargeMocks. At higher \(k\), the EuclidLargeMocks tend to overestimate the quadrupole and, in the last redshift bin, slightly underestimate the monopole by about **7%**. The 2PCF multipoles likewise show excellent agreement, and any apparent BAO peak shift is consistent with sample variance. A key technical result is that EuclidLargeMocks do not suffer from the replication-induced cross-redshift-bin correlations that affect Geppetto [2507.12116].

The same paper also tests cosmological parameter inference using an **EFT-based model**, the Gaussian-process emulator **Comet**, and the sampler **NAUTILUS**. The fitted data vector comprises the power-spectrum monopole and quadrupole in **four redshift bins**, with cosmological parameters
\[
\{h,\omega_{\rm c},A_{\rm s},n_{\rm s},\omega_{\rm b}\}
\]
and nuisance parameters
\[
\{b_1,b_2,\gamma_{21}\}
\]
per redshift bin. The priors include wide uniform priors on most cosmological parameters, together with
\[
\mathcal{N}(0.96,0.041)
\]
for \(n_s\) and
\[
\mathcal{N}(2.218\times10^{-2},\,0.055\times10^{-2})
\]
for \(\omega_{\rm b}\). Using either the Flagship mock as data or one EuclidLargeMocks realization as data yields consistent posteriors within \(\sim1\sigma\), both for conservative cuts \(k<0.2\,h\) and more aggressive cuts \(k<0.3\,h\). This establishes that one EuclidLargeMocks realization reproduces the Flagship posterior within sample variance [2507.12116].

The principal caveat in this validation chain is the already identified quadrupole excess at \(k>0.2\,h\,\mathrm{Mpc}^{-1}\), which the paper links to the HOD/mass calibration on one-halo scales. The effect is reported not to bias cosmological inference at the tested scales, but it bounds how aggressively the mocks can be pushed.

## 4. Real-space theory validation with Flagship I H\(\alpha\) mocks

A complementary branch of the Euclid mock programme benchmarks real-space galaxy power-spectrum models against very large H\(\alpha\)-selected catalogues extracted from **four comoving snapshots** of the Euclid **Flagship I** \(N\)-body simulation at
\[
z=(0.9,1.2,1.5,1.8).
\]
Flagship I evolves
\[
2\times10^{12}
\]
particles in a box of side
\[
L=3780\,h^{-1}\mathrm{Mpc},
\]
corresponding to a comoving volume of about
\[
58\,h^{-3}\,\mathrm{Gpc}^3.
\]
The snapshots are populated with H\(\alpha\) emitters using halo occupation distributions tuned to the Flagship light-cone catalogue, specifically the Euclid H\(\alpha\) Model 1 and Model 3 prescriptions from Pozzetti et al. The resulting catalogues contain millions of galaxies overall; the table in the paper lists counts ranging from roughly \(1.1\times10^8\) at \(z=0.9\) down to \(\sim1.7\times10^7\) at \(z=1.8\), with mean number densities from \(\sim3.7\times10^{-3}\) to \(\sim3\times10^{-4}\,(h/\mathrm{Mpc})^3\). These samples are intentionally optimistic, assuming
\[
f_{\mathrm{H}\alpha}=2\times10^{-16}\,\mathrm{erg\,cm^{-2}\,s^{-1}},
\]
no interlopers, and no observational incompleteness or survey-mask effects [2312.00679].

The paper compares two model classes. The first is a third-order Eulerian EFTofLSS bias expansion:
\[
\delta_g(x)=b_1\,\delta(x)+b_2\,\frac{\delta^2(x)}{2}+\cdots +\nabla^2\delta(x)+\varepsilon_g(x),
\]
including the tidal operators \(\mathcal{G}_2\) and \(\Gamma_3\). The galaxy power spectrum is decomposed as
\[
P_{gg}(k)=P_{gg}^{\rm tree}(k)+P_{gg}^{\rm 1-loop}(k)+P_{gg}^{\rm ctr}(k)+P_{gg}^{\rm noise}(k),
\]
with
\[
P_{gg}^{\rm tree}(k)=b_1^2 P_{mm}(k),
\qquad
P_{gg}^{\rm ctr}(k)=-2\,c_0\,k^2 P_{mm}(k),
\]
and non-Poissonian shot noise
\[
P_{gg}^{\rm noise}(k)=\frac{1}{\bar n}(1+P_{,\!1}+P_{,\!2}k^2).
\]
The most general Eulerian nuisance set contains six free parameters:
\[
b_1,\ b_2,\ b_{\mathcal G_2},\ b_{\Gamma_3},\ c_0,\ P_{,\!1}.
\]
The second model is a hybrid Lagrangian perturbation theory plus high-resolution simulation approach implemented through the **BACCO emulator**, with operator basis
\[
\{1,\delta,\delta^2,s^2,\nabla^2\delta\},
\]
Lagrangian biases \(b_1^{\mathcal L}\), \(b_2^{\mathcal L}\), \(b_{s^2}^{\mathcal L}\), \(b_{\nabla^2\delta}^{\mathcal L}\), and stochastic amplitude \(P_{,\!1}\) [2312.00679].

The headline validation result is that both models remain unbiased in the \((h,\omega_c)\) plane up to at least
\[
k_{\max}\simeq 0.45\,h\,\mathrm{Mpc}^{-1}
\]
for all four redshifts when the covariance is rescaled to Euclid-like shell volumes. For the EFTofLSS description, the preferred configuration is to fix the quadratic tidal bias to the excursion-set relation
\[
b_{\mathcal G_2}^{\rm ex-set}=0.524-0.547\,b_1+0.046\,b_1^2,
\]
while optionally fixing the cubic tidal bias to the coevolution relation
\[
b_{\mathcal G_2}^{\rm coev}=-\frac{2}{7}(b_1-1),\qquad
b_{\Gamma_3}^{\rm coev}=-\frac{1}{6}(b_1-1)-\frac{5}{2}b_{\mathcal G_2}^{\rm L}.
\]
The BACCO model is reported as competitive and unbiased only when a **\(0.5\%\)** theory error is included; without it, emulator imperfections can induce noticeable bias, especially at low redshift and high \(k\). In Euclid-like shells over \(15{,}000\,\mathrm{deg}^2\), with \(\Delta z=(0.2,0.2,0.2,0.3)\) and covariance rescaling
\[
C_{\rm shell}=\eta\,C_{\rm box},
\qquad
\eta=\frac{V_{\rm box}}{V_{\rm shell}},
\]
with \(\eta\) ranging from about 3.3 to 6.7, the same conclusions hold [2312.00679].

These tests are not the DR1 light-cone EuclidLargeMocks themselves, but they form a theory-validation counterpart for the same spectroscopic target population and redshift range.

## 5. Redshift interlopers and contaminated DR1 catalogues

The interloper extension of EuclidLargeMocks is constructed to quantify catastrophic redshift errors in the Euclid spectroscopic sample. The contaminated sample is partitioned into **correct galaxies**, **line interlopers**, and **noise interlopers**. For wrong-line identification, the redshift mapping is written as
\[
1+z=\lambda_{\rm obs}/\lambda_{\rm rest},
\]
and
\[
\frac{1+z_{\rm true}}{1+z_{\rm meas}}=\frac{\lambda_{\rm wrong}}{\lambda_{\rm true}}.
\]
The main contaminant lines identified are \(\mathrm{OIII}\,5008\) and \(\mathrm{SIII}\,9531\), with representative mean fractions in four spectroscopic bins:
\(z\in[0.9,1.1]\): OIII 0.03, SIII 0.01, noise 0.12;
\(z\in[1.1,1.3]\): OIII 0.12, SIII 0.03, noise 0.08;
\(z\in[1.3,1.5]\): OIII 0.09, SIII 0.08, noise 0.08;
\(z\in[1.5,1.8]\): OIII 0.01, SIII 0.07, noise 0.06 [2505.04688].

The contaminated density field is decomposed as
\[
\delta(\mathbf{x}) = (1-f_{\rm tot})\,\delta_{\rm c}(\mathbf{x})
+ \sum_i f_i\,\delta_i(\mathbf{x}_{\parallel}/\gamma_{\parallel}, \mathbf{x}_{\perp}/\gamma_{\perp})
+ f_{\rm n}\int\!\!\int \mathcal{P}_{\rm n}(\gamma_{\parallel},\gamma_{\perp}) \,\delta_{\rm n}(\mathbf{x}_{\parallel}/\gamma_{\parallel}, \mathbf{x}_{\perp}/\gamma_{\perp}) \,d\gamma_{\parallel}\,d\gamma_{\perp},
\]
with geometric remapping factors
\[
\gamma_{\perp}=\frac{D_A(z_{\rm meas})}{D_A(z_{\rm true})},
\qquad
\gamma_{\parallel}=\frac{(1+z_{\rm meas})/H(z_{\rm meas})}{(1+z_{\rm true})/H(z_{\rm true})}.
\]
The 2PCF is measured with the **Landy–Szalay estimator**, and the mocks allow direct measurement of all component terms: correct-correct, line-line, noise-noise, correct-line, correct-noise, and line-noise. The results show that the correct-galaxy term dominates everywhere; line interlopers are subdominant but non-negligible, especially around \(z\sim1.3-1.5\); correct-noise cross-correlation is important at low redshift; and all other cross-terms are negligible. A central conclusion is that the contaminated 2PCF is not just a rescaled version of the uncontaminated one, because line interlopers can shift and broaden the BAO feature [2505.04688].

The modelling hierarchy culminates in a minimal attenuation-only description in which the correct-galaxy clustering term is multiplied by
\[
p_c=(1-f_{\rm tot})^2.
\]
For Euclid DR1, this minimal model is sufficient to recover the correct values of \(f\sigma_8\), \(\alpha_{\parallel}\), and \(\alpha_{\perp}\). The induced systematic error on \(f\sigma_8\) is about **1%–3%**, depending on redshift, and remains smaller than the expected DR1 statistical error. The AP parameters are largely insensitive to interlopers: contaminated and uncontaminated posteriors are nearly identical, and even the worst-case systematic shift is well below the expected DR1 statistical error and below DR3-level statistical errors. The paper therefore concludes that DR1 full-shape analyses may model interlopers primarily as a loss of clustering amplitude, whereas more detailed treatments will become more important for future, higher-precision releases and for smaller, more nonlinear scales [2505.04688].

## 6. Relativistic projection effects and related covariance methodology

A separate EuclidLargeMocks-style light-cone programme targets **relativistic redshift-space distortions** on the largest scales of the Euclid Wide Spectroscopic Survey. It consists of **140 mock galaxy catalogues** generated from **35 independent simulations** of side
\[
L_{\rm box}=12\,h^{-1}\mathrm{Gpc},
\]
each yielding **four non-overlapping light cones**. The mocks use the **LIGER** method in large-box mode, which maps Newtonian simulation outputs to observed redshift space to linear order in perturbations. The EWSS galaxy population is specified through \(\bar n_g(z)\), \(b(z)\), \(\mathcal E(z)\), and \(\mathcal Q(z)\), with
\[
F_{\rm lim}=2\times10^{-16}\,\mathrm{erg\,cm^{-2}\,s^{-1}}
\]
and a **70% completeness factor**, and linear bias
\[
b(z)=1.46+0.68\,(z-1).
\]
The survey mask removes \(20^\circ\) around the Galactic and ecliptic planes, leaving **four disconnected sky patches** [2410.00956].

For each light cone, four variants are generated: **R** (real space), **V** (velocity-only RSD), **G** (velocity plus integrated terms, especially lensing, but no observer velocity), and **O** (all effects, including observer velocity). The linear observed overdensity is written schematically as
\[
\delta_{g,s} = \delta^{\rm com}_g -\frac{1}{H}\frac{\partial(\mathbf v_e\cdot \mathbf n)}{\partial x} +2(\mathcal Q-1)\kappa +\cdots,
\]
and in the mock implementation as
\[
\delta_{g,s}=(b-1)\delta_m^{\rm com}+\delta_s+\mathcal E\,\delta\ln a+\mathcal Q(\mathcal M-1).
\]
The measured summary statistics are \(C_\ell\), \(\xi_\ell(r)\), and \(P_\ell(k)\). Weak-lensing magnification and convergence emerge as the dominant relativistic correction: their signal-to-noise ranges from **2.5 to 6**, depending on the statistic, with the broad bin \(0.9<z<1.8\) reaching **S/N \(\approx 5.4\)** in the angular analysis; the highest redshift bin reaches **S/N \(\approx 2.5\)** in \(\xi_\ell\) and **S/N \(\approx 2.4\)** in \(P_\ell(k)\). By contrast, the imprint of the observer velocity is modest, with **S/N < 1** in 2PCF multipoles and only about **1.0–1.3** in power-spectrum multipoles. When all relativistic effects are included, the window-corrected Kaiser model that keeps only velocity-gradient RSD is rejected for \(1.5<z<1.8\) at **2.9\(\sigma\)**. The paper also shows that the mixing-matrix formalism for finite-volume effects remains robust for the EWSS’s disconnected survey geometry [2410.00956].

A related methodological strand in Euclid preparation addresses covariance calibration rather than galaxy mock construction. For the real-space 2PCF of galaxy clusters, **1000 PINOCCHIO light cones** over **10,313 deg\(^2\)** and \(z=0\) to \(2.5\) were used to validate a semi-analytical covariance model. The baseline Gaussian plus Poisson shot-noise model underestimates diagonal terms by about **30%** and off-diagonal terms by about **50%** at intermediate and high redshift. Introducing fitted parameters \(\alpha\), \(\beta\), and \(\gamma\) yields covariance accuracy at the **10%** level and reduces figure-of-merit differences to about **5%** for \((\Omega_{\rm m},\sigma_8)\); the cosmology-dependent covariance is statistically preferred, with
\[
\langle \Delta{\rm DIC}\rangle_{\rm sims} = -11.5\pm1.6.
\]
This cluster result is not a EuclidLargeMocks galaxy-catalogue result, but it illustrates the broader Euclid reliance on large mock ensembles for covariance validation [2211.12965].

## 7. Flagship 2, SciPICal, and the evolution of Euclid mock production

A later development connects EuclidLargeMocks to **Euclid Flagship 2** through **SciPIC**, a modular halo-to-galaxy population pipeline, and **SciPICal**, an automated calibration layer designed to optimize the mock properties that most affect clustering. In this framework, the halo catalogue inputs include halo mass, comoving position and velocity, halo shape or ellipsoid, and concentration; the outputs include number of galaxies per halo, luminosities, colours, SEDs, and later derived properties. SciPICal tunes the parameters controlling halo occupation, central- and satellite-galaxy luminosities, colours, and satellite positions, and is implemented on **Apache Spark** on the **PIC Big Data platform** using Hadoop, with **600 CPU cores** for mock generation and **480 CPU cores** for clustering measurements [2604.14802].

For **FS2-Wide**, the underlying halo catalogue has box side **3600 \(h^{-1}\) Mpc**, particle mass **\(10^9\,h^{-1} M_\odot\)**, softening length **\(4.5\,h^{-1}\) kpc**, a full-sky particle light-cone to **\(z=3\)**, halo finder **ROCKSTAR**, and total size **126 billion** main haloes. The calibration subset is a **\(300\,h^{-1}\) Mpc** sub-volume containing **8,075,637 haloes**. For **FS2-Deep**, the box side is **1000 \(h^{-1}\) Mpc**, the particle mass is **\(10^8\,h^{-1} M_\odot\)**, and the product includes **two opposite-sky light-cones**, each **50 deg\(^2\)**, extending to **\(z=10\)**, together with **100 redshift snapshots** and a complementary \(z=0\) snapshot; the deep-survey total area is stated as about **53 deg\(^2\)** [2604.14802].

The six calibrated parameters are
\[
\Theta=\{\alpha, f, \sigma_L, M_r^\mathrm{red}, M_r^\mathrm{green}, f^\mathrm{cut}_r\},
\]
with initial bounds
\[
\alpha\in[0.8,1.2],\quad
f\in[13,20],\quad
\sigma_L\in[0.05,0.3],\quad
M_r^\mathrm{red}\in[-23,-15],\quad
M_r^\mathrm{green}\in[-23,-15],\quad
f^\mathrm{cut}_r\in[0.5,3.5].
\]
The satellite occupation is modeled as
\[
\langle N_\mathrm{sats} \rangle = \left( \frac{M_\mathrm{h}}{f \, M_\mathrm{min}} \right)^\alpha,
\]
and the main calibration objective is a projected-clustering \(\chi^2\),
\[
\chi^2 = \sum_i \frac{\left[w_{\mathrm{p},i}(\Theta)-w^\mathrm{obs}_{\mathrm{p},i}\right]^2}{\sigma^2_{2\mathrm{p},i}}.
\]
Against SDSS DR7 / Zehavi et al. (2011), the calibrated FS2-Wide version improves the reduced \(\chi^2\) from approximately **4** to approximately **2**, i.e. about **50%**. The paper reports good agreement, within **15% for most of the samples**, across validations against spectroscopic and photometric surveys and a hydrodynamical simulation. It also states that the largest remaining limitations are in reproducing **colour-selected clustering**, especially for low-brightness red galaxies, and that peculiar velocities are ignored during calibration even though they are used later for validation [2604.14802].

This suggests a longer-term architectural shift in the Euclid mock programme: DR1-oriented Pinocchio/HOD catalogues provide the immediate covariance-ready and pipeline-ready ensemble, while Flagship 2 plus SciPICal provides an update path for wide and deep mocks whose galaxy properties can be recalibrated as new observational constraints become available.

Source: https://www.emergentmind.com/topics/euclidlargemocks