Quaia Data Set: Gaia-unWISE Quasar Catalog
- Quaia is a Gaia-unWISE quasar catalog combining Gaia DR3 candidates and unWISE photometry with proper-motion and color-based cleaning for enhanced redshift accuracy.
- It offers two magnitude-limited samples (G<20.5 and G<20.0) over an all-sky coverage, with specific selection-function products enabling robust cosmological analysis.
- Its derivative products—mock catalogs, void/cluster catalogs, and radio cross-matches—support studies on ultra-large-scale structure, growth history, and primordial non-Gaussianity.
Searching arXiv for papers on the Quaia data set and related analyses. Quaia is the Gaia-unWISE Quasar Catalog, an all-sky quasar data set assembled by combining Gaia DR3 quasar candidates with unWISE infrared photometry, followed by proper-motion and color-based cleaning and machine-learning redshift refinement. In its released form, it provides a magnitude-limited sample of 1,295,502 quasars at and a cleaner subset of 755,850 objects at , together with selection-function products designed for cosmological analysis (Storey-Fisher et al., 2023). In subsequent work, Quaia has also become the parent data set for mock catalogs, void and cluster catalogs, radio cross-matches, and field-level reconstructions of large-scale structure, making it both a survey catalog and a platform for derivative data products (Sinigaglia et al., 19 Sep 2025, Arsenov et al., 22 Sep 2025, Andrews et al., 2 Feb 2026).
1. Catalog genesis and construction
Quaia originates from the Gaia DR3 quasar-candidate sample and is refined by cross-matching to unWISE, whose mid-infrared photometry is used to improve purity and redshift estimation. The core selection pipeline applies cuts based on proper motions and Gaia and unWISE colors, reducing the number of contaminants by . Redshifts are then improved with a -nearest-neighbors model trained on SDSS quasars. For the sample, this procedure yields only catastrophic errors with and catastrophic errors with , corresponding to reductions of and 0, respectively, relative to the Gaia redshifts (Storey-Fisher et al., 2023).
The catalog is routinely described as a spectro-photometric or hybrid spectro-photometric quasar sample. That designation is consequential: Quaia is not a uniform spectroscopic survey in the narrow sense used for fiber-based redshift programs, but a large-area quasar catalog whose redshift information combines Gaia low-resolution spectra, unWISE photometry, and SDSS-trained machine learning (Storey-Fisher et al., 2023, Andrews et al., 2 Feb 2026). This combination is what allows Quaia to sample extremely large comoving volumes while remaining usable for clustering, cross-correlation, and population studies.
A persistent strength of Quaia is that it was designed together with explicit selection-function products. Subsequent analyses repeatedly rely on those products, rather than treating the catalog as a raw source list, which is essential for any interpretation of angular density fluctuations on ultra-large scales (Alonso et al., 2024, Fabbian et al., 29 Apr 2025).
2. Sample definitions, coverage, and redshift characterization
Quaia is usually analyzed through two principal magnitude-limited samples. In field-level reconstruction work these are labeled Quaia Deep (1, 1,295,502 quasars, median 2) and Quaia Clean (3, 755,850 quasars, median 4); both are described there as spanning 5 over 6, corresponding to 7 after conservative cuts (Andrews et al., 2 Feb 2026). Other analyses adopt stricter masks, so the effective footprint is study-dependent rather than unique to the catalog.
| Sample or subset | Definition | Size / area |
|---|---|---|
| Quaia Deep | 8 | 1,295,502 sources |
| Quaia Clean | 9 | 755,850 sources |
| GW-association subset | 0, 1 | 660,031 objects |
| High-2 cosmic-web subset | 3, selection function 4 | 708,483 quasars over 5 |
The catalog’s “all-sky” characterization therefore coexists with extensive masking practice. In projected analyses, maps are pixelized with HEALPix and pixels with selection function 6 are often masked; in the void-and-cluster analysis, a more conservative threshold of selection function 7 leaves 8 and 708,483 quasars in the range 9 (Alonso et al., 2024, Arsenov et al., 22 Sep 2025). In gravitational-wave host-environment studies, the main geometric cut is 0, motivated by Galactic contamination and extinction, leaving 660,031 objects at 1 (Veronesi et al., 2024).
Redshift characterization is similarly context-dependent. In the catalog-centered description summarized for the mock-catalog paper, the final 2 sample has 3 of sources with 4 and 5 with 6 relative to SDSS (Sinigaglia et al., 19 Sep 2025). In projected large-scale-structure analyses, a typical working characterization is 7 for the 8 sample (Fabbian et al., 29 Apr 2025). For completeness-sensitive applications, these redshift errors are supplemented by explicit completeness models in redshift and luminosity space. One such analysis computed completeness in 8 redshift bins from 9 to 0 and for 5 bolometric-luminosity thresholds, finding, for example, that at 1 completeness is 2 for 3 and remains above 4 for 5 (Veronesi et al., 2024).
3. Selection functions, map making, and systematic control
A defining methodological feature of Quaia is that analyses treat the angular selection function as part of the data set. In the main projected-clustering studies, the selection function 6 is constructed using Gaussian process regression to model mean density fluctuations induced by observational systematics, including Galactic dust and scanning-related effects (Alonso et al., 2024). In the original catalog description, the selection is explicitly conceptualized as a product of Gaia, unWISE, and post-selection terms,
7
with 8, 9, and 0 denoting position, magnitude, and color (Storey-Fisher et al., 2023).
Map-level analyses commonly construct quasar overdensity fields as
1
where 2 is the object count in a pixel, 3 is the pixelized selection function, and 4 is the mean expected count (Villagra et al., 11 Jul 2025, Arcari et al., 26 Sep 2025). This correction is fundamental in Quaia-based analyses because the dominant cosmological applications exploit very low multipoles and very large angular scales, where even small spatially varying systematics can bias inference.
Tomographic organization is another recurring feature. Several projected studies split Quaia at the median redshift 5, producing low-6 and high-7 bins for clustering and CMB-lensing cross-correlations (Fabbian et al., 29 Apr 2025, Villagra et al., 11 Jul 2025). The growth-history analysis instead used three bins centered at 8 (Piccirilli et al., 2024), whereas BORG-based field-level inference used eight radial bins with approximately equal comoving volumes (Andrews et al., 2 Feb 2026). Redshift distributions are frequently estimated by stacking individual Gaussian redshift PDFs, rather than by histogramming point estimates alone (Piccirilli et al., 2024, Villagra et al., 11 Jul 2025).
The literature also makes clear that systematic mitigation is not optional. Templates for dust extinction, stellar contamination, scanning patterns, Magellanic-Cloud regions, and large-scale dipoles are explicitly modeled, masked, or projected out. In primordial non-Gaussianity analyses, systematic modes are removed by mode deprojection in both pseudo-9 and QML pipelines, with transfer-function corrections calibrated on simulations (Fabbian et al., 29 Apr 2025). In dipole analyses, progressively stricter Galactic masks are needed to suppress residual contamination near the Galactic plane and Galactic center (Mittal et al., 2023).
4. Cosmological inference from Quaia
Quaia’s main cosmological role has been as a tracer of ultra-large-scale structure at redshifts that are difficult to access with lower-0 galaxy surveys. A representative example is the measurement of growth history from quasar clustering and its cross-correlation with CMB lensing. Using three tomographic bins centered at 1, 2, and 3, one analysis obtained 4, 5, and 6, with the highest-redshift point described as one of the highest-7 measurements of 8 made with that technique (Piccirilli et al., 2024).
A second line of work uses Quaia to probe the turnover of the matter power spectrum through the quasar auto-spectrum and its cross-correlation with Planck CMB lensing. In that analysis, the turnover is detected with a significance between 9 and 0, depending on method, and the scale parameter is measured as 1, corresponding to 2 precision on the equality scale (Alonso et al., 2024). That measurement was then used, in combination with distance-ladder information, to infer 3, 4, and 5 under the assumptions described in the study (Alonso et al., 2024).
Quaia has also become a leading projected-data set for constraints on local-type primordial non-Gaussianity. Using quasar auto-correlations and cross-correlations with Planck PR4 CMB lensing in two redshift bins, one study obtained 6 at 7 confidence for the universality-response choice 8, and 9 from lensing cross-correlations alone (Fabbian et al., 29 Apr 2025). A later analysis introduced angular redshift fluctuations (ARF) as an additional projected observable and reported 0, described there as a 1 improvement over the previous Quaia measurement and the tightest result achieved with two-point projected summary statistics (Bermejo-Climent et al., 23 Jan 2026).
High-redshift growth constraints have been extended with ACT DR6 and Planck PR4 lensing. In a joint 2pt analysis using the Quaia auto-spectrum, quasar–lensing cross-spectra, the ACT lensing auto-spectrum, and BOSS BAO, the inferred value is 3, while the reconstructed fluctuation amplitude at the median redshift of the high-4 signal is 5 (Villagra et al., 11 Jul 2025). This suggests that Quaia’s utility is not confined to its own median redshift range: through the broad CMB-lensing kernel, it also helps constrain growth beyond 6.
5. Synthetic catalogs and derivative data products
The Quaia ecosystem includes a substantial mock-catalog program. The dedicated mock-catalog paper presents 100 full-sky spectrophotometric quasar mocks with smooth redshift evolution from 7 to 8, built from dark-matter light cones generated with WebON using Eulerian Augmented Lagrangian Perturbation Theory over a 9 volume on a 0 grid. Quasar biasing is modeled with the Hicobian hierarchical nonlocal nonlinear bias scheme, calibrated to AbacusSummit HOD catalogs tuned to DESI EDR observations, after which the mocks are degraded with spectro-photometric redshift uncertainties, the Quaia angular selection function, and the observed 1 (Sinigaglia et al., 19 Sep 2025). The post-processing includes the redshift perturbation
2
and the resulting catalogs are validated against full-sky maps, redshift-error distributions, 3, angular power spectra, normalized covariance matrices, and angular two-point correlation functions, with the paper reporting excellent agreement (Sinigaglia et al., 19 Sep 2025).
Quaia has also been used to construct catalogs of cosmic voids and clusters at high redshift. Using 708,483 quasars in 4 over 5, the REVOLVER/ZOBOV Voronoi-tessellation pipeline reconstructs local densities via
6
and identifies 12,842 voids and 41,111 clusters (Arsenov et al., 22 Sep 2025). The largest structures reach 7 for voids and 8 for clusters, while the agreement between data and 50 mocks is reported at the 9 level for radii, average inner density, and density profiles (Arsenov et al., 22 Sep 2025). The authors additionally release value-added catalogs linking individual quasars to local density, corrected Voronoi volume, and void/cluster membership.
A distinct derivative product is the Quaia–VLASS catalog, formed by cross-matching Quaia with the VLASS 3 GHz radio catalog. Using a 00 arcsec matching radius, that work produces 43,650 cross-matched sources and finds a radio-loud fraction of 01 for the full sample, with no significant large-scale pattern in radio loudness across the sky (Arsenov et al., 2024). The value-added records include extinction-corrected 02-band magnitudes, radio flux densities, radio magnitudes, completeness values, and quality flags.
At the field level, BORG has been applied directly to Quaia Deep and Quaia Clean to infer initial conditions, present-day dark matter, and velocity fields over a 03 volume with maximum spatial resolution 04. That study describes the reconstruction as the largest field-level reconstruction of the observable Universe by comoving volume to date and reports a 05 cross-correlation with Planck CMB lensing as an external validation (Andrews et al., 2 Feb 2026).
6. Scientific interpretation, limitations, and contested points
Quaia’s scale and uniformity have made it attractive for tests of isotropy, but those tests also illustrate the catalog’s limitations. In the Bayesian analysis of the cosmic dipole, the raw sample exhibits residual selection effects and contamination near the Galactic plane. After excising increasingly conservative latitude bands, especially at 06, the lower-contamination sample is found to be consistent in amplitude and direction with the CMB dipole, whereas less aggressively masked data prefer more complex dipolar structure attributed to contamination rather than cosmology (Mittal et al., 2023). A common misconception is therefore that the catalog’s all-sky character alone guarantees isotropy-grade cleanliness; the published analyses do not support that view.
The same pattern appears in other cosmological uses. Primordial-non-Gaussianity work masks low-selection pixels, deprojects systematic templates, and excludes the lowest multipoles in the quasar auto-spectrum (Fabbian et al., 29 Apr 2025). Turnover and growth analyses likewise adopt scale cuts and cross-correlation strategies specifically because the quasar auto-spectrum is more vulnerable to observational artifacts than its CMB-lensing cross-spectrum (Alonso et al., 2024, Villagra et al., 11 Jul 2025). This suggests that the main scientific value of Quaia lies not only in source count and sky area, but in the combination of source count, explicit survey modeling, and a workflow that treats systematics as first-class inferential objects.
Beyond large-scale-structure cosmology, Quaia has been used in multi-messenger and fundamental-physics contexts. In gravitational-wave host-population studies, the combination of near-all-sky coverage and completeness estimates as functions of redshift and luminosity allows statistical tests of AGN-origin scenarios for black-hole mergers; one analysis concludes at 95 per cent credibility that unobscured AGN above 07 and 08 do not contribute to more than 21 per cent and 11 per cent, respectively, of detected GW events (Veronesi et al., 2024). In CMB-polarization cross-correlation work, the first measurement of the cross-spectrum between anisotropic birefringence from Planck NPIPE and Quaia galaxy counts is consistent with the null hypothesis, with PTE 09 and fitted scale-invariant amplitude 10, yet still yields new constraints on axion-photon coupling in the ultra-light regime (Arcari et al., 26 Sep 2025).
Taken together, the published record presents Quaia as a high-volume, high-redshift quasar data set whose scientific reach depends critically on accompanying metadata: selection functions, masks, redshift-error models, tomographic binning, and increasingly realistic mocks. It is best understood not as a single static catalog, but as a structured survey framework supporting projected clustering, CMB cross-correlations, field-level inference, environmental reconstruction, and value-added multiwavelength catalogs across a broad range of cosmological and astrophysical applications (Storey-Fisher et al., 2023, Sinigaglia et al., 19 Sep 2025).