Papers
Topics
Authors
Recent
Search
2000 character limit reached

CatNorth Database: Photometric Quasar Catalogue

Updated 12 July 2026
  • CatNorth Database is a photometric quasar-candidate catalogue assembled using Gaia DR3, Pan-STARRS1, and CatWISE2020, featuring over 1.5 million high-purity sources with both photometric and spectroscopic redshifts.
  • It employs ensemble classification methods including XGBoost and Random Forest, achieving ~90% purity and enabling precise quasar selection for cosmological studies and strong-lensing searches.
  • The database supports diverse applications such as S8 cosmological parameter inference and lensed quasar discovery, and delivers detailed data products like astrometry, redshift estimates, and spatial selection function maps.

CatNorth is a photometric quasar-candidate database assembled from Gaia DR3, Pan-STARRS1, and CatWISE2020, designed for quasar science, large-scale-structure analyses, cosmological parameter inference, and strong-lensing searches. The catalogue was introduced as an improved Gaia DR3 quasar candidate resource with more than 1.5 million sources in the 3π3\pi sky, high purity, photometric redshifts for all candidates, and spectroscopic redshifts for a substantial subset from Gaia BP/RP spectra (Fu et al., 2023). Subsequent work uses CatNorth both as a homogeneous parent sample for wide-separation lensed-quasar searches (Wu et al., 21 Sep 2025) and as the basis of a machine-learning-corrected quasar clustering data set for S8S_8 measurements with Planck DR4 CMB lensing (Qin et al., 10 Mar 2026).

1. Survey basis, scope, and sample definition

CatNorth is built from three survey layers. The parent source list is the Gaia DR3 qso_candidates table, comprising 6.6\sim 6.6 million objects to G21G\approx 21. These Gaia candidates are cross-matched to Pan-STARRS1 3π\pi photometry in g,r,i,z,yg,r,i,z,y and to CatWISE2020 mid-infrared photometry in W1,W2W1,W2 (Fu et al., 2023). In the cosmology-oriented description, the database is characterized as a photometric quasar-candidate catalogue specifically assembled for large-scale-structure and cosmological studies, with high purity, accurate redshifts, and a well-characterized spatial selection function over a very large area of sky (Qin et al., 10 Mar 2026).

The sky footprint is approximately the PS1 3π3\pi survey region, with declination δ30\delta \gtrsim -30^\circ (Fu et al., 2023). For cosmological analyses, the nominal coverage is the full “3π3\pi” sky with S8S_80, excluding a small southern patch for simplicity, and with a Galactic-plane-related mask implemented through the completeness threshold S8S_81 (Qin et al., 10 Mar 2026). The wide-separation lensing study describes the footprint as S8S_82 steradians, while emphasizing that regions poorly covered by Pan-STARRS or WISE, including the Galactic plane and southernmost declinations, are not used (Wu et al., 21 Sep 2025).

The catalogue is reported in two closely related magnitude regimes. The original release contains S8S_83 quasar candidates at S8S_84, while a brighter subset contains S8S_85 candidates at S8S_86 (Fu et al., 2023). In the cosmological application, the masked S8S_87 sample contains S8S_88 million sources over S8S_89 after excising low-completeness regions (Qin et al., 10 Mar 2026). Across these uses, the stated purity is 6.6\sim 6.60 (Fu et al., 2023, Wu et al., 21 Sep 2025).

2. Classification architecture and selection logic

The catalogue release specifies an ensemble classification pipeline based on XGBoost. Two base classifiers, CLF_LVAC and CLF_GDR3, are trained using different master stellar samples, and each outputs 6.6\sim 6.61, 6.6\sim 6.62, and 6.6\sim 6.63. The ensemble quasar probability is defined by

6.6\sim 6.64

with analogous ensemble probabilities for stars and galaxies (Fu et al., 2023).

The training design is explicitly heterogeneous. Extragalactic training objects include 6.6\sim 6.65 SDSS DR16Q quasars with high-quality redshifts and 6.6\sim 6.66 SDSS DR17 spectroscopic galaxies without broad lines. The stellar side uses two “master” samples, each 6.6\sim 6.67 million objects, augmented by ultracool dwarfs, white dwarfs, and carbon stars (Fu et al., 2023). Input features are the 14 colors 6.6\sim 6.68, 6.6\sim 6.69, G21G\approx 210, G21G\approx 211, G21G\approx 212, G21G\approx 213, G21G\approx 214, G21G\approx 215, G21G\approx 216, G21G\approx 217, G21G\approx 218, G21G\approx 219, and π\pi0, together with the corrected BP/RP flux-excess factor π\pi1. Proper motions and parallax enter indirectly through the zero-proper-motion probability density π\pi2 (Fu et al., 2023).

The operational selection cuts in the release paper are a high-quasar-likelihood threshold,

π\pi3

and a proper-motion consistency cut,

π\pi4

On held-out validation, the classifier attains balanced-accuracy π\pi5, weighted π\pi6, and MCC π\pi7. The proper-motion cut removes π\pi8 of stars while retaining π\pi9 of quasars (Fu et al., 2023).

A notable interpretive point is that the wide-separation lensed-quasar paper summarizes CatNorth as using a Random Forest classifier based on parallax significance, proper-motion significance, and multi-band colors, with stellar-like g,r,i,z,yg,r,i,z,y0 or g,r,i,z,yg,r,i,z,y1 removed (Wu et al., 21 Sep 2025). This indicates that descriptions of the classification backend differ across papers. The common operational outcome is a Gaia–PS1–CatWISE quasar-candidate sample with estimated purity near g,r,i,z,yg,r,i,z,y2 and photometric redshifts for all retained objects (Fu et al., 2023, Wu et al., 21 Sep 2025).

3. Redshift inference, data fields, and delivered products

CatNorth provides photometric redshifts for all candidates through an ensemble regression model. The training sample comprises g,r,i,z,yg,r,i,z,y3 DR16Q quasars plus g,r,i,z,yg,r,i,z,y4 Milliquas quasars at g,r,i,z,yg,r,i,z,y5 or g,r,i,z,yg,r,i,z,y6, for a total of g,r,i,z,yg,r,i,z,y7 objects. Fifteen inputs are used: the 14 colors and Gaia’s lower and upper redshift confidence limits transformed as g,r,i,z,yg,r,i,z,y8 and g,r,i,z,yg,r,i,z,y9. The three component regressors are XGBoost, TabNet, and FT-Transformer, ensembled by averaging (Fu et al., 2023).

On a validation set of W1,W2W1,W20 quasars, the photometric-redshift ensemble achieves W1,W2W1,W21, W1,W2W1,W22, and outlier fraction W1,W2W1,W23, where

W1,W2W1,W24

and

W1,W2W1,W25

In the cosmology paper, the practical redshift range is summarized as W1,W2W1,W26 to W1,W2W1,W27, with typical uncertainties W1,W2W1,W28–W1,W2W1,W29 (Fu et al., 2023, Qin et al., 10 Mar 2026).

For a subset of the catalogue, Gaia BP/RP spectra are used to infer spectroscopic redshifts with a convolutional neural network. The release paper describes a RegNet CNN taking calibrated BP+RP spectra sampled from 3π3\pi0–3π3\pi1 Å at 3π3\pi2 Å resolution, for 300 input pixels, with four repeated Conv1D–ReLU–MaxPool blocks followed by fully connected layers. It reports validation metrics 3π3\pi3, 3π3\pi4, and 3π3\pi5 on 3π3\pi6 quasars (Fu et al., 2023). The cosmology summary states that spectroscopic redshifts are available for 3π3\pi7 objects from Gaia BP/RP spectra using a convolutional neural network (Qin et al., 10 Mar 2026).

At the data-model level, CatNorth entries include astrometry, Gaia photometry, PS1 PSF magnitudes and errors, CatWISE2020 3π3\pi8 magnitudes and errors, class probabilities, and redshift products (Fu et al., 2023). The lensing paper adds summary astrometric significances,

3π3\pi9

and the color-similarity statistic δ30\delta \gtrsim -30^\circ0 used for grouped-image comparison (Wu et al., 21 Sep 2025). The cosmology paper describes distribution through FITS or HDF5 catalogue tables, HEALPix maps for the selection function δ30\delta \gtrsim -30^\circ1, overdensity δ30\delta \gtrsim -30^\circ2, and mask, query access via standard VO protocols or a public GitHub/GitLab repository, and analysis notebooks using astropy, healpy/HEALPix, pyccl, pymaster/NaMaster, emcee, numpy, scipy, matplotlib, and PyTorch (Qin et al., 10 Mar 2026).

4. Selection function formalism and angular-systematics control

For cosmological applications, CatNorth is accompanied by an explicit spatial selection-function formalism. The observed catalogue is modeled as

δ30\delta \gtrsim -30^\circ3

Restricting to angular systematics,

δ30\delta \gtrsim -30^\circ4

and the overdensity field is defined as

δ30\delta \gtrsim -30^\circ5

Pixels with δ30\delta \gtrsim -30^\circ6 are masked to avoid numerical instabilities (Qin et al., 10 Mar 2026).

The selection function δ30\delta \gtrsim -30^\circ7 is estimated with a neural network using 10 normalized systematics templates: δ30\delta \gtrsim -30^\circ8, a Gaia scanning-law proxy δ30\delta \gtrsim -30^\circ9, 3π3\pi0stellar density3π3\pi1, five PS1 median-magnitude maps in 3π3\pi2, and two CatWISE2020 median-magnitude maps in 3π3\pi3. The architecture is

3π3\pi4

with ReLU activations and Dropout regularization (Qin et al., 10 Mar 2026).

Training is restricted to “clean” pixels defined as the top 3π3\pi5 least-extincted or deepest regions in each template. The quantity 3π3\pi6 is normalized to unity in those regions, and the loss is the mean-squared error between predicted 3π3\pi7 and observed 3π3\pi8 in clean pixels (Qin et al., 10 Mar 2026). Reported stability is high: the pixel-to-pixel variation in 3π3\pi9 is S8S_800 over 30 repeated trainings, and after correction the cross-power spectra between S8S_801 and each systematics template are consistent with zero. Mock simulations with an injected selection function are reported to confirm no over-suppression of large-scale power (Qin et al., 10 Mar 2026).

The delivered map products reflect this analysis design. The selection function is stored as a HEALPix FITS map at S8S_802, while the overdensity field is computed at S8S_803 by up-sampling S8S_804 and dividing the pixel counts (Qin et al., 10 Mar 2026). For power-spectrum work, the recommended cuts are S8S_805 to avoid cosmic-variance bias and Limber-breakdown effects, and S8S_806 to remain in the linear regime and within NaMaster reliability. Broad redshift bins, including a two-bin split at S8S_807, are explicitly recommended to preserve selection-function fidelity and reduce photo-S8S_808 leakage (Qin et al., 10 Mar 2026).

5. Derived subsamples and scientific use cases

CatNorth has been used in at least three distinct modes: as a quasar-population catalogue, as a parent sample for lensed-quasar discovery, and as a cosmological tracer sample. The release paper states that it is the main source of input catalog for the LAMOST phase III quasar survey, which is expected to build a highly complete sample of bright quasars with S8S_809 (Fu et al., 2023).

In strong-lensing work, CatNorth serves as the parent sample for a HEALPix-based friends-of-friends search for wide-separation lensed quasars. All sources are assigned to HEALPix pixels with S8S_810, corresponding to angular resolution S8S_811. Grouping proceeds through isolated multi-object pixels and FoF chaining across adjacent pixels, followed by filters based on intra-group color and spectral similarity. The search considers separations between S8S_812 and S8S_813 arcsec and reduces the S8S_814 sources to S8S_815 groups while retaining all known, discoverable WSLQs (Wu et al., 21 Sep 2025). The resulting candidate list contains 333 new WSLQ candidates with separations from S8S_816 to S8S_817 arcsec. Using SDSS DR16 and DESI DR1 spectroscopy, two new candidate systems are uncovered; the remaining 331 candidates lack sufficient spectra and are labeled as 45 grade A, 98 grade B, and 188 grade C. A by-product sample of 29 confirmed dual quasars is also compiled (Wu et al., 21 Sep 2025).

For cosmology, CatNorth is partitioned into flux-limited and volume-limited subsamples. The flux-limited S8S_818 split contains a S8S_819 bin with 574,411 sources before masking and 518,037 after S8S_820, and a S8S_821 bin with 574,410 before masking and 515,532 after. The effective redshifts are S8S_822 and S8S_823, respectively. The volume-limited samples are S8S_824; S8S_825; and S8S_826, with masked counts 771,827, 479,424, and 469,686 (Qin et al., 10 Mar 2026).

Subsample Selection Reported S8S_827
Flux-limited low-S8S_828 S8S_829 S8S_830
Flux-limited high-S8S_831 S8S_832 S8S_833
Volume-limited S8S_834 S8S_835
Volume-limited S8S_836 S8S_837
Volume-limited S8S_838 S8S_839

These measurements are compared in the cosmology paper to the Planck 2018 CMB anisotropy constraint S8S_840 and to a previously reported value S8S_841 from the Quaia quasar candidate catalog. The stated conclusion is that current CatNorth-based measurements show less evidence of the S8S_842 tension (Qin et al., 10 Mar 2026).

The catalogue has several explicit limitations. The magnitude limit S8S_843 excludes fainter quasar images, which directly affects strong-lensing completeness; among 8 published wide-separation lensed quasars, only 4 are “discoverable” in CatNorth, meaning they have at least two counterpart images with S8S_844 (Wu et al., 21 Sep 2025). The footprint is limited by PS1 and WISE coverage and excludes the Galactic plane and southernmost declinations (Wu et al., 21 Sep 2025, Fu et al., 2023).

Photometric-redshift uncertainties remain relevant even after the ensemble regression design. The cosmology paper notes that emission-line misidentification, especially C IV versus C III], is largely reduced but leaves residual S8S_845 scatter S8S_846, and it recommends marginalizing over S8S_847 uncertainties through redshift shifts S8S_848 and width changes S8S_849; in the reported tests, no significant S8S_850 shift is found (Qin et al., 10 Mar 2026). It also states that narrow redshift bins are disfavored because they reduce the number of objects per pixel, amplify shot noise, and increase sensitivity of S8S_851–S8S_852 modeling to photo-S8S_853 errors (Qin et al., 10 Mar 2026).

The angular selection function corrects multiplicative biases such as depth variations and stellar contamination, but it does not remove additive errors, including unmodeled contamination spikes. Best practice therefore includes masking known contaminants such as M31/M33 and the LMC/SMC, together with all pixels with S8S_854 (Qin et al., 10 Mar 2026). For CMB lensing cross-correlations, the paper cautions that high-S8S_855 measurements can be biased by Cosmic Infrared Background contamination of the Planck DR4 S8S_856-map; the lower S8S_857 values at S8S_858 are interpreted as likely reflecting residual incompleteness and/or foreground bias (Qin et al., 10 Mar 2026).

Bias modeling is another controlled source of uncertainty in the cosmology application. The fiducial model

S8S_859

is reported to fit well. Tests with constant, steep, and three-parameter free S8S_860 give consistent low-S8S_861 S8S_862, while high-S8S_863 S8S_864 is more sensitive though the main conclusions are unchanged (Qin et al., 10 Mar 2026).

A further practical caveat is that CatNorth is not strictly volume- or flux-limited as a target-selection function. The wide-separation lensing study notes that the classifier trades completeness for purity and that the target-selection function can be mapped through the training-set confusion matrix (Wu et al., 21 Sep 2025). This suggests that CatNorth is best interpreted not as a single statistically complete quasar census, but as a high-purity, multi-purpose quasar-candidate infrastructure whose effective selection must be modeled differently in population studies, strong-lens searches, and clustering-based cosmology.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to CatNorth Database.