CatNorth Database: Photometric Quasar Catalogue
- CatNorth Database is a photometric quasar-candidate catalogue assembled using Gaia DR3, Pan-STARRS1, and CatWISE2020, featuring over 1.5 million high-purity sources with both photometric and spectroscopic redshifts.
- It employs ensemble classification methods including XGBoost and Random Forest, achieving ~90% purity and enabling precise quasar selection for cosmological studies and strong-lensing searches.
- The database supports diverse applications such as S8 cosmological parameter inference and lensed quasar discovery, and delivers detailed data products like astrometry, redshift estimates, and spatial selection function maps.
CatNorth is a photometric quasar-candidate database assembled from Gaia DR3, Pan-STARRS1, and CatWISE2020, designed for quasar science, large-scale-structure analyses, cosmological parameter inference, and strong-lensing searches. The catalogue was introduced as an improved Gaia DR3 quasar candidate resource with more than 1.5 million sources in the sky, high purity, photometric redshifts for all candidates, and spectroscopic redshifts for a substantial subset from Gaia BP/RP spectra (Fu et al., 2023). Subsequent work uses CatNorth both as a homogeneous parent sample for wide-separation lensed-quasar searches (Wu et al., 21 Sep 2025) and as the basis of a machine-learning-corrected quasar clustering data set for measurements with Planck DR4 CMB lensing (Qin et al., 10 Mar 2026).
1. Survey basis, scope, and sample definition
CatNorth is built from three survey layers. The parent source list is the Gaia DR3 qso_candidates table, comprising million objects to . These Gaia candidates are cross-matched to Pan-STARRS1 3 photometry in and to CatWISE2020 mid-infrared photometry in (Fu et al., 2023). In the cosmology-oriented description, the database is characterized as a photometric quasar-candidate catalogue specifically assembled for large-scale-structure and cosmological studies, with high purity, accurate redshifts, and a well-characterized spatial selection function over a very large area of sky (Qin et al., 10 Mar 2026).
The sky footprint is approximately the PS1 survey region, with declination (Fu et al., 2023). For cosmological analyses, the nominal coverage is the full “” sky with 0, excluding a small southern patch for simplicity, and with a Galactic-plane-related mask implemented through the completeness threshold 1 (Qin et al., 10 Mar 2026). The wide-separation lensing study describes the footprint as 2 steradians, while emphasizing that regions poorly covered by Pan-STARRS or WISE, including the Galactic plane and southernmost declinations, are not used (Wu et al., 21 Sep 2025).
The catalogue is reported in two closely related magnitude regimes. The original release contains 3 quasar candidates at 4, while a brighter subset contains 5 candidates at 6 (Fu et al., 2023). In the cosmological application, the masked 7 sample contains 8 million sources over 9 after excising low-completeness regions (Qin et al., 10 Mar 2026). Across these uses, the stated purity is 0 (Fu et al., 2023, Wu et al., 21 Sep 2025).
2. Classification architecture and selection logic
The catalogue release specifies an ensemble classification pipeline based on XGBoost. Two base classifiers, CLF_LVAC and CLF_GDR3, are trained using different master stellar samples, and each outputs 1, 2, and 3. The ensemble quasar probability is defined by
4
with analogous ensemble probabilities for stars and galaxies (Fu et al., 2023).
The training design is explicitly heterogeneous. Extragalactic training objects include 5 SDSS DR16Q quasars with high-quality redshifts and 6 SDSS DR17 spectroscopic galaxies without broad lines. The stellar side uses two “master” samples, each 7 million objects, augmented by ultracool dwarfs, white dwarfs, and carbon stars (Fu et al., 2023). Input features are the 14 colors 8, 9, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, and 0, together with the corrected BP/RP flux-excess factor 1. Proper motions and parallax enter indirectly through the zero-proper-motion probability density 2 (Fu et al., 2023).
The operational selection cuts in the release paper are a high-quasar-likelihood threshold,
3
and a proper-motion consistency cut,
4
On held-out validation, the classifier attains balanced-accuracy 5, weighted 6, and MCC 7. The proper-motion cut removes 8 of stars while retaining 9 of quasars (Fu et al., 2023).
A notable interpretive point is that the wide-separation lensed-quasar paper summarizes CatNorth as using a Random Forest classifier based on parallax significance, proper-motion significance, and multi-band colors, with stellar-like 0 or 1 removed (Wu et al., 21 Sep 2025). This indicates that descriptions of the classification backend differ across papers. The common operational outcome is a Gaia–PS1–CatWISE quasar-candidate sample with estimated purity near 2 and photometric redshifts for all retained objects (Fu et al., 2023, Wu et al., 21 Sep 2025).
3. Redshift inference, data fields, and delivered products
CatNorth provides photometric redshifts for all candidates through an ensemble regression model. The training sample comprises 3 DR16Q quasars plus 4 Milliquas quasars at 5 or 6, for a total of 7 objects. Fifteen inputs are used: the 14 colors and Gaia’s lower and upper redshift confidence limits transformed as 8 and 9. The three component regressors are XGBoost, TabNet, and FT-Transformer, ensembled by averaging (Fu et al., 2023).
On a validation set of 0 quasars, the photometric-redshift ensemble achieves 1, 2, and outlier fraction 3, where
4
and
5
In the cosmology paper, the practical redshift range is summarized as 6 to 7, with typical uncertainties 8–9 (Fu et al., 2023, Qin et al., 10 Mar 2026).
For a subset of the catalogue, Gaia BP/RP spectra are used to infer spectroscopic redshifts with a convolutional neural network. The release paper describes a RegNet CNN taking calibrated BP+RP spectra sampled from 0–1 Å at 2 Å resolution, for 300 input pixels, with four repeated Conv1D–ReLU–MaxPool blocks followed by fully connected layers. It reports validation metrics 3, 4, and 5 on 6 quasars (Fu et al., 2023). The cosmology summary states that spectroscopic redshifts are available for 7 objects from Gaia BP/RP spectra using a convolutional neural network (Qin et al., 10 Mar 2026).
At the data-model level, CatNorth entries include astrometry, Gaia photometry, PS1 PSF magnitudes and errors, CatWISE2020 8 magnitudes and errors, class probabilities, and redshift products (Fu et al., 2023). The lensing paper adds summary astrometric significances,
9
and the color-similarity statistic 0 used for grouped-image comparison (Wu et al., 21 Sep 2025). The cosmology paper describes distribution through FITS or HDF5 catalogue tables, HEALPix maps for the selection function 1, overdensity 2, and mask, query access via standard VO protocols or a public GitHub/GitLab repository, and analysis notebooks using astropy, healpy/HEALPix, pyccl, pymaster/NaMaster, emcee, numpy, scipy, matplotlib, and PyTorch (Qin et al., 10 Mar 2026).
4. Selection function formalism and angular-systematics control
For cosmological applications, CatNorth is accompanied by an explicit spatial selection-function formalism. The observed catalogue is modeled as
3
Restricting to angular systematics,
4
and the overdensity field is defined as
5
Pixels with 6 are masked to avoid numerical instabilities (Qin et al., 10 Mar 2026).
The selection function 7 is estimated with a neural network using 10 normalized systematics templates: 8, a Gaia scanning-law proxy 9, 0stellar density1, five PS1 median-magnitude maps in 2, and two CatWISE2020 median-magnitude maps in 3. The architecture is
4
with ReLU activations and Dropout regularization (Qin et al., 10 Mar 2026).
Training is restricted to “clean” pixels defined as the top 5 least-extincted or deepest regions in each template. The quantity 6 is normalized to unity in those regions, and the loss is the mean-squared error between predicted 7 and observed 8 in clean pixels (Qin et al., 10 Mar 2026). Reported stability is high: the pixel-to-pixel variation in 9 is 00 over 30 repeated trainings, and after correction the cross-power spectra between 01 and each systematics template are consistent with zero. Mock simulations with an injected selection function are reported to confirm no over-suppression of large-scale power (Qin et al., 10 Mar 2026).
The delivered map products reflect this analysis design. The selection function is stored as a HEALPix FITS map at 02, while the overdensity field is computed at 03 by up-sampling 04 and dividing the pixel counts (Qin et al., 10 Mar 2026). For power-spectrum work, the recommended cuts are 05 to avoid cosmic-variance bias and Limber-breakdown effects, and 06 to remain in the linear regime and within NaMaster reliability. Broad redshift bins, including a two-bin split at 07, are explicitly recommended to preserve selection-function fidelity and reduce photo-08 leakage (Qin et al., 10 Mar 2026).
5. Derived subsamples and scientific use cases
CatNorth has been used in at least three distinct modes: as a quasar-population catalogue, as a parent sample for lensed-quasar discovery, and as a cosmological tracer sample. The release paper states that it is the main source of input catalog for the LAMOST phase III quasar survey, which is expected to build a highly complete sample of bright quasars with 09 (Fu et al., 2023).
In strong-lensing work, CatNorth serves as the parent sample for a HEALPix-based friends-of-friends search for wide-separation lensed quasars. All sources are assigned to HEALPix pixels with 10, corresponding to angular resolution 11. Grouping proceeds through isolated multi-object pixels and FoF chaining across adjacent pixels, followed by filters based on intra-group color and spectral similarity. The search considers separations between 12 and 13 arcsec and reduces the 14 sources to 15 groups while retaining all known, discoverable WSLQs (Wu et al., 21 Sep 2025). The resulting candidate list contains 333 new WSLQ candidates with separations from 16 to 17 arcsec. Using SDSS DR16 and DESI DR1 spectroscopy, two new candidate systems are uncovered; the remaining 331 candidates lack sufficient spectra and are labeled as 45 grade A, 98 grade B, and 188 grade C. A by-product sample of 29 confirmed dual quasars is also compiled (Wu et al., 21 Sep 2025).
For cosmology, CatNorth is partitioned into flux-limited and volume-limited subsamples. The flux-limited 18 split contains a 19 bin with 574,411 sources before masking and 518,037 after 20, and a 21 bin with 574,410 before masking and 515,532 after. The effective redshifts are 22 and 23, respectively. The volume-limited samples are 24; 25; and 26, with masked counts 771,827, 479,424, and 469,686 (Qin et al., 10 Mar 2026).
| Subsample | Selection | Reported 27 |
|---|---|---|
| Flux-limited low-28 | 29 | 30 |
| Flux-limited high-31 | 32 | 33 |
| Volume-limited | 34 | 35 |
| Volume-limited | 36 | 37 |
| Volume-limited | 38 | 39 |
These measurements are compared in the cosmology paper to the Planck 2018 CMB anisotropy constraint 40 and to a previously reported value 41 from the Quaia quasar candidate catalog. The stated conclusion is that current CatNorth-based measurements show less evidence of the 42 tension (Qin et al., 10 Mar 2026).
6. Limitations, ambiguities, and recommended practice
The catalogue has several explicit limitations. The magnitude limit 43 excludes fainter quasar images, which directly affects strong-lensing completeness; among 8 published wide-separation lensed quasars, only 4 are “discoverable” in CatNorth, meaning they have at least two counterpart images with 44 (Wu et al., 21 Sep 2025). The footprint is limited by PS1 and WISE coverage and excludes the Galactic plane and southernmost declinations (Wu et al., 21 Sep 2025, Fu et al., 2023).
Photometric-redshift uncertainties remain relevant even after the ensemble regression design. The cosmology paper notes that emission-line misidentification, especially C IV versus C III], is largely reduced but leaves residual 45 scatter 46, and it recommends marginalizing over 47 uncertainties through redshift shifts 48 and width changes 49; in the reported tests, no significant 50 shift is found (Qin et al., 10 Mar 2026). It also states that narrow redshift bins are disfavored because they reduce the number of objects per pixel, amplify shot noise, and increase sensitivity of 51–52 modeling to photo-53 errors (Qin et al., 10 Mar 2026).
The angular selection function corrects multiplicative biases such as depth variations and stellar contamination, but it does not remove additive errors, including unmodeled contamination spikes. Best practice therefore includes masking known contaminants such as M31/M33 and the LMC/SMC, together with all pixels with 54 (Qin et al., 10 Mar 2026). For CMB lensing cross-correlations, the paper cautions that high-55 measurements can be biased by Cosmic Infrared Background contamination of the Planck DR4 56-map; the lower 57 values at 58 are interpreted as likely reflecting residual incompleteness and/or foreground bias (Qin et al., 10 Mar 2026).
Bias modeling is another controlled source of uncertainty in the cosmology application. The fiducial model
59
is reported to fit well. Tests with constant, steep, and three-parameter free 60 give consistent low-61 62, while high-63 64 is more sensitive though the main conclusions are unchanged (Qin et al., 10 Mar 2026).
A further practical caveat is that CatNorth is not strictly volume- or flux-limited as a target-selection function. The wide-separation lensing study notes that the classifier trades completeness for purity and that the target-selection function can be mapped through the training-set confusion matrix (Wu et al., 21 Sep 2025). This suggests that CatNorth is best interpreted not as a single statistically complete quasar census, but as a high-purity, multi-purpose quasar-candidate infrastructure whose effective selection must be modeled differently in population studies, strong-lens searches, and clustering-based cosmology.