---
title: Quasar Boosted Model Approaches
url: https://www.emergentmind.com/topics/quasar-boosted-model
type: topic
---

# Quasar Boosted Model Approaches

to=arxiv_search.search  红鼎json code
{"query":"quasar boosted model photometric variability damped random walk 1009.2081 boosted random forest quasar detection 1304.0401 XGBoost quasars J-PLUS 2212.05868 CatBoost quasar photometric redshift 2605.09450", "max_results": 10}
to=arxiv_search.search  天天买彩票json code
{"query":"1009.2081 Quasar Selection Based on Photometric Variability", "max_results": 5}
to=arxiv.search 񎔇json code
{"query":"1009.2081 Quasar Selection Based on Photometric Variability", "max_results": 5}
Quasar Boosted Model is not a single standardized formalism. In the literature, the phrase is most naturally applied to quasar methods in which an added statistical descriptor or an explicit boosting layer materially improves discrimination or inference. The clearest early usage is the variability-selection framework of "Quasar Selection Based on Photometric Variability," where the characteristic damped-random-walk time scale $\tau$ is added to variability-amplitude information and measurably improves quasar selection relative to structure-function-slope methods [1009.2081]. In later survey work, the same phrase can reasonably denote AdaBoosted Random Forest, XGBoost, and CatBoost-based quasar classifiers or regressors operating on light-curve or multiband photometric features [1304.0401], [1903.03335], [2212.05868], [2605.09450].

## 1. Terminological scope

In the variability-selection literature, the “boost” is not an ensemble of weak learners but the increase in selection performance obtained by including the DRW damping time scale $\tau$ in addition to variability amplitude. In survey machine learning, by contrast, “boosted” has its standard algorithmic meaning: sequential ensemble learning, as in AdaBoost, XGBoost, or CatBoost. The term therefore spans at least two distinct methodological lineages: physically motivated stochastic variability modeling and boosted decision-tree classification or regression [1009.2081], [1304.0401], [1903.03335], [2212.05868], [2605.09450].

This suggests that “Quasar Boosted Model” functions more as a context-dependent label than as a canonical model name. In one context it denotes a quasar selector whose discriminative power is boosted by time-scale information; in another it denotes boosted-tree systems for catalog classification, quasar-candidate ranking, or photometric-redshift estimation. The distinction matters, because later quasar models that are stronger or more probabilistic are not necessarily boosting-based.

## 2. Damped-random-walk variability selection

The archetypal form of the concept is the SDSS Stripe 82 variability selector developed for separating quasars from other variable point sources. The method uses the $g$-band light curves of the Stripe 82 variable point-source catalog, requiring at least ten observations, rms variability in $g$ and $r$ exceeding $0.05$ mag, and $\chi^2/{\rm dof}>3$ for a constant-flux fit. The full variable-object sample with $i<19$ contains $52{,}547$ objects, among which $1{,}912$ ($4\%$) are spectroscopically confirmed quasars. For the main extragalactic, lower-contamination test, the restriction $-35^{\circ}<{\rm RA}<50^{\circ}$ leaves $10{,}024$ variable sources, including $1{,}490$ ($15\%$) confirmed quasars; the spectroscopic quasar sample is explicitly noted as complete for $i<19$ in the quasar color region [1009.2081].

Quasar variability is modeled as a damped random walk, equivalently an Ornstein–Uhlenbeck process, with exponential covariance
$$
S_{ij}=\sigma^2\exp(-|t_i-t_j|/\tau).
$$
Here $\tau$ is the damping time scale and $\sigma$ is the long-term standard deviation of the variability. The short-timescale driving amplitude is defined as
$$
\hat{\sigma}=\sigma\sqrt{2/\tau}.
$$
In structure-function form,
$$
SF(\Delta t)={\rm SF}_{\infty}\left(1-e^{-|\Delta t|/\tau}\right)^{1/2},
$$
with
$$
SF(\Delta t\gg\tau)\equiv {\rm SF}_{\infty}=\hat{\sigma}\sqrt{\tau},
$$
and, for short lags,
$$
SF(\Delta t\ll\tau)={\rm SF}_{\infty}\sqrt{\frac{|\Delta t|}{\tau}}=\hat{\sigma}\sqrt{|\Delta t|}.
$$
Accordingly, ${\rm SF}_{\infty}$ is the asymptotic rms variability amplitude, while $\tau$ controls where the structure function flattens.

The comparative baseline is the older structure-function-slope parameterization used in Schmidt et al. (2010),
$$
SF(\Delta t)=A(\Delta t/1\,{\rm yr})^\gamma,
$$
which effectively traces short-timescale behavior but does not uniquely recover $\tau$. The DRW fit instead yields individual $\tau$ and $\hat{\sigma}$ values for each light curve. The operative claim of the model is that quasars are not merely variable; they are variable on characteristic time scales that differ from many stellar contaminants. In this formulation, the “boost” comes from recovering that time-scale information explicitly rather than compressing variability to a slope-like summary [1009.2081].

## 3. Selection metrics, thresholds, and survey-scale implications

Within the Stripe 82 lower-contamination sample, completeness and efficiency are defined as
$$
C=\frac{\#~{\rm of~selected~confirmed~quasars}}{\#~{\rm of~confirmed~quasars}}\times 100,
\qquad
E=\frac{\#~{\rm of~selected~confirmed~quasars}}{\#~{\rm of~selected~objects}}\times 100.
$$
Using variability alone, inclusion of $\tau$ boosts efficiency from about $60\%$ to $75\%$ while maintaining $C=98\%$. For fixed completeness $C=90\%$, efficiency improves from about $80\%$ to $85\%$. Conversely, if efficiency is held at $E=80\%$, completeness improves from $90\%$ to $97\%$ once $\tau$ is included. The paper also reports that selecting quasars with both $\tau$ and $\hat{\sigma}$ and without color information can achieve $E=82\%$ with $C=96\%$, or $E=75\%$ with $C=98\%$, and $E=85\%$ with $C=90\%$ [1009.2081].

Threshold behavior is central to the operational model. The paper adopts $\tau\ge 100$ days as the optimal single-parameter cut because completeness remains high while efficiency is close to its asymptotic value. For this simple cut, the reported result is $C=94\%$ and $E=81\%$. If the most outlying point in each light curve is rejected and $\Delta L_{\rm noise}>2$ is required, efficiency rises to $E=87\%$ while completeness drops slightly to $C=93\%$. A stricter $\Delta L_{\rm noise}>10$ cut with $\tau\ge 10^{1.5}$ days yields $C=96\%$ and $E=87\%$. With quasar-like colors added on top of the variability cut, purity increases further: for $\tau\ge 100$ days and quasar colors in regions II and IV, $C=91\%$ and $E=96\%$. For the UV-excess subset, $\tau\ge 100$ days alone gives $C=95\%$ and $E=97\%$, while for the non-UV-excess subset it gives $C=100\%$ and $E=69\%$, the latter being limited by spectroscopic incompleteness [1009.2081].

The same framework is explicitly projected to future synoptic surveys. For a simulated LSST cadence over $10$ years with photometric accuracy $0.03$ mag at $i\approx 22$, a simple $\tau>100$ day criterion is expected to give $C=88\%$; the same threshold yields $C=75\%$ for a $3$-year light curve and $C=51\%$ for a $1$-year light curve, demonstrating the importance of long baselines for reliable $\tau$ recovery. For Pan-STARRS1 $3\pi$, with roughly $24$ combined $griz$ epochs over $3$ years, the estimated completeness is $C\approx 70\%$ for $\tau>100$ days. The DRW inference scales linearly with the number of data points, $O(N)$, which is relevant for LSST-scale photometric archives. The paper further argues that, given adequate survey cadence, photometric variability can outperform color selection in some regimes, especially near $z\sim 3$ and for reddened or otherwise non-standard quasars missed by color cuts [1009.2081].

## 4. AdaBoosted Random Forest light-curve classifiers

A second, algorithmically distinct meaning of Quasar Boosted Model appears in the EROS-2 and MACHO light-curve literature. "An improved quasar detection method in EROS-2 and MACHO LMC datasets" uses a two-stage ensemble in which Random Forest is the base learner and AdaBoost is the sequential boosting wrapper, producing the boosted Random Forest classifier denoted AB+RF [1304.0401].

Each light curve is represented by $14$ features per band, or $28$ features total in the two-band surveys. Eleven are earlier time-series features, and three new features per band are derived from a continuous auto-regressive model. The CAR(1) process is written as
$$
dX(t)=-\frac{1}{\tau}X(t)\,dt+\sigma_C\sqrt{dt}\,\epsilon(t)+b\,dt,
\qquad \tau,\sigma_C,t\ge 0,
$$
with mean
$$
\mathbb{E}[X(t)]=b\tau
$$
and variance
$$
{\rm Var}(X(t))=\frac{\tau\sigma_C^2}{2}.
$$
The paper argues that quasars tend to have large $\tau$, and that $\sigma_C$ helps separate quasars from many non-variable stars and some periodic variables. For computational speed on tens of millions of objects, it estimates only $(\sigma_C,\tau)$ directly and computes $b$ as the mean magnitude divided by $\tau$; the reduced chi-square difference relative to full three-parameter optimization is reported as less than $2.5\%$ on average.

The training sets are survey-specific. The EROS-2 training set contains $65$ known quasars, $67$ Be stars, $330$ long-period stars, $5829$ non-variable stars, $1727$ RR Lyrae, $406$ Cepheids, and $488$ eclipsing binaries. The MACHO training set contains $3969$ non-variable stars, $127$ Be stars, $78$ Cepheids, $193$ eclipsing binaries, $288$ RR Lyrae, $574$ microlensing events, $359$ long-period variables, and $58$ quasars. Under $10$-fold cross-validation, the reported F-scores for AB+RF with CAR features are $0.868$ on EROS-2 and $0.877$ on MACHO, exceeding the corresponding SVM and plain Random Forest results. The abstract summarizes the training-set performance as about $90\%$ precision and $86\%$ recall. Applied to the full databases, the model identifies $1160$ EROS-2 candidates and $2551$ MACHO candidates. The paper also notes that about $25\%$ of false positives are periodic stars, implying that a dedicated periodic-star filter could further improve the classifier [1304.0401].

## 5. Gradient-boosted photometric classification and photometric redshifting

In later survey work, the boosted formulation is primarily a tabular-data method operating on multiband photometry, colors, morphology, extinction, and survey-specific metadata. The following representative systems illustrate the transition from variability-based boosting to catalog-scale boosted-tree inference.

| Paper | Task and boosted model | Key reported result |
|---|---|---|
| "Efficient Selection of Quasar Candidates Based on Optical and Infrared Photometric Data Using Machine Learning" [1903.03335] | Pan-STARRS1 + AllWISE star–quasar classification with XGBoost on the 8Color feature set | Accuracy $99.46\%$ with default parameters and $99.58\%$ after hyperparameter optimization; $2{,}006{,}632$ intersected sources with $P_{\rm QSO}>0.5$, of which $1{,}201{,}211$ have $P_{\rm QSO}>0.95$ |
| "J-PLUS DR3: Galaxy-Star-Quasar classification" [2212.05868] | TPOT-selected XGBoost for three-class galaxy–star–quasar classification using $37$ J-PLUS features | AUC above $0.99$ for galaxies, stars, and quasars; AP above $0.99$ for galaxies and stars and above $0.96$ for quasars; value-added catalog for $47{,}431{,}242$ sources |
| "Search for quasar pairs with Gaia astrometric data. II. Photometric redshift prediction with machine learning for the MGQPC catalogue" [2605.09450] | CatBoost photometric-redshift point estimation plus FlexZBoost redshift-PDF estimation for quasar-pair triage | Normalised median absolute deviation $0.036$ and outlier fraction $5.6\%$ on the test sample; application to MGQPC yields $185$ high-probability quasar-pair candidates, including $20$ spectroscopically confirmed physical pairs |

The Pan-STARRS1–AllWISE system is explicitly motivated by the observation that combined optical and infrared colors are more effective than optical colors alone for star–quasar separation. Its XGBoost classifier is a boosted ensemble of regression trees with prediction
$$
\hat{y}_i=\sum_{k=1}^{K}f_k(x_i),
$$
regularized objective
$$
Obj=\sum_i l(\hat{y}_i,y_i)+\sum \Omega(f_k),
$$
and the binary-logistic objective for classification. The best-performing input is the 8Color set
$$
g-r,\; r-i,\; i-z,\; z-y,\; y-W1,\; W1-W2,\; i-W1,\; z-W2.
$$
The paper also constructs the $yW1W2$ and $iW1zW2$ color cuts, but the central methodological point is that boosted trees can learn nonlinear multicolor boundaries more effectively than manually designed two-dimensional separators [1903.03335].

The J-PLUS DR3 study extends the boosted paradigm to three-class classification. It trains on a spectroscopic crossmatch with SDSS DR18, LAMOST DR8, and Gaia, yielding a final training set of $2{,}092{,}930$ sources split into $664{,}687$ galaxies, $1{,}161{,}203$ stars, and $267{,}040$ quasars. TPOT searches over pipelines using ROC AUC as the optimization metric, with generations $=100$, population size $=100$, and offspring size $=100$, for a total of $10{,}100$ analyzed pipelines; the selected model is XGBoost with the original $37$ features and no transformations or stacking. The paper emphasizes that quasar classification remains the hardest class because quasars are rarer than stars, are often point-like, and the training set is less representative at faint magnitudes [2212.05868].

The MGQPC work uses boosted trees for a different purpose: photometric-redshift inference rather than direct quasar–nonquasar discrimination. CatBoost supplies point estimates, while FlexZBoost provides full redshift PDFs through conditional density estimation. The normalized residual is defined as
$$
\Delta z=\frac{z_{\rm spec}-z_{\rm phot}}{1+z_{\rm spec}},
$$
with
$$
\sigma_{\rm NMAD}=1.4826\times {\rm median}(|\Delta z|),
$$
and outlier fraction
$$
\eta=\frac{N(|\Delta z|>0.15)}{N_{\rm total}}.
$$
Pair candidates are filtered using the line-of-sight velocity difference
$$
|\Delta v(z_A,z_B)|=c\cdot \frac{|z_A-z_B|}{1+(z_A+z_B)/2},
$$
with the fiducial criterion $|\Delta v|<2000~{\rm km\,s^{-1}}$. This is a boosted model in the operational sense of high-throughput catalog triage, where point estimates and calibrated PDFs are used together to prioritize rare physical quasar pairs for spectroscopy [2605.09450].

## 6. Distinct but frequently confused formulations

Several important quasar models are stronger or more physically structured than earlier approaches, yet are explicitly not boosting-based. Quasar Factor Analysis models the observed spectrum as
$$
S=C\circ \exp(-\tau(z))+\omega(z)+\epsilon,
$$
with the continuum
$$
C=\mu+Fh+\Psi,
$$
and learns the parameters by maximizing the average log-likelihood over spectra. Both arXiv versions explicitly describe QFA as a probabilistic unsupervised latent-factor model rather than a boosting-based method; its function is continuum prediction, posterior inference, and outlier detection, not ensemble boosting [2207.02788], [2211.11784].

A separate, non-ML use of the idea appears in quasar demographics. "Comparing Simple Quasar Demographics Models" constructs the observed quasar luminosity function by combining a self-consistent black-hole/galaxy history with an intrinsic luminosity tied either to $M_{\rm BH}$ or $\dot{M}_{\rm BH}$, and then applying a stochastic variability distribution, either lognormal or truncated power law. In that setting, the observed luminosity function is “boosted” or scattered relative to the intrinsic one by variability, rather than by an ensemble-learning algorithm [1409.2832].

Another distinct usage occurs in models of quasar-induced Ly$\alpha$ emission from the intergalactic medium. There, the “quasar-boosted” picture refers to a foreground quasar enhancing surrounding Ly$\alpha$ emission through resonant scattering and, more importantly, fluorescence from optically thick absorbers. The predicted quasar-induced contribution accounts for only about $10\%$ of the BOSS/eBOSS measurements in the outer region $>10\ {\rm cMpc}\ h^{-1}$, and the paper concludes that the quasar-induced component alone is not sufficient, though it is not in conflict with the data once star-forming galaxies are included [2312.02374].

Taken together, these literatures indicate that Quasar Boosted Model is best interpreted contextually. In time-domain quasar selection, it most precisely denotes the DRW-based classifier whose time-scale parameter $\tau$ boosts efficiency and completeness. In survey machine learning, it denotes boosted ensembles such as AdaBoosted Random Forest, XGBoost, and CatBoost. In adjacent quasar theory and spectroscopy, superficially similar language may instead refer to variability convolution, latent-factor generative modeling, or quasar-induced radiative enhancement rather than boosting in the ensemble-learning sense.

Source: https://www.emergentmind.com/topics/quasar-boosted-model