Papers
Topics
Authors
Recent
Search
2000 character limit reached

Scan Statistics for Nonhomogeneous Poisson Processes with Extreme-Value Calibration and Application to CNV Detection

Published 11 Jun 2026 in stat.ME and stat.CO | (2606.13973v1)

Abstract: We develop a scan statistic method for detecting local clusters in a two-sample nonhomogeneous Poisson process (NHPP) framework, motivated by copy number variation (CNV) analysis in next-generation sequencing data. The control sample is used to construct an empirical time transformation, under which the transformed case sample is approximately uniform on [0,1] under the null hypothesis. The scan statistic is defined as the maximum number of transformed points within a moving window. We show that the scan statistic converges to a generalized extreme value (GEV) distribution with an extremal index that captures the dependence induced by overlapping windows. The GEV parameters and extremal index are estimated using maximum likelihood and exceedance clustering methods, providing an asymptotic calibration of the test. A permutation procedure is also developed to provide a nonparametric alternative. Simulation studies show that the permutation calibration maintains empirical Type I error close to the nominal level across the considered settings, and the GEV calibration is accurate for smaller windows. Both proposed procedures show competitive power compared with the continuous testing method under heterogeneous baseline intensities. An application to sequencing data illustrates the effectiveness of the proposed approach for detecting CNV regions.

Authors (2)

Summary

  • The paper develops a two-sample nonhomogeneous Poisson process scan statistic that uses an empirical time transformation and an extremal-index-adjusted Gumbel limit to calibrate localized cluster detection without modeling the heterogeneous baseline parametrically.
  • Permutation calibration maintains Type I error near the nominal 5% level across simulated baseline shapes, while GEV calibration is competitive for shrinking windows but can become slightly liberal when windows are large and overlap dependence is stronger.
  • In synthetic sequencing data, both proposed methods consistently recover three main CNV regions across window sizes, matching SeqCBS closely and generally outperforming CTS for small-to-moderate windows, although CTS can be more powerful under rapidly increasing baselines with large windows.

Problem and motivation

This paper develops a scan statistic framework for detecting localized clusters in a two-sample nonhomogeneous Poisson process (NHPP) setting, motivated by copy number variation (CNV) and copy number alteration (CNA) detection in next-generation sequencing (NGS) data. The control (normal) and case (tumor) read locations are modeled as independent NHPPs on [0,1][0,1], with the null hypothesis that λ1(t)=ρλ0(t)\lambda_1(t)=\rho\lambda_0(t) for some ρ>0\rho>0. Under the alternative, an unknown interval of length ss exhibits locally elevated (amplification) or depressed (deletion) case intensity. The central inferential difficulty is calibrating the tail of the scan statistic—the maximum count over sliding windows—when the baseline intensity is heterogeneous, unknown, and induces strong dependence among overlapping windows. Existing rigorous null distributions for scan statistics under NHPP models are limited, particularly in the two-sample setting; homogeneous-baseline approximations inflate Type I error because large local counts may reflect baseline heterogeneity rather than a true cluster.

The work is positioned relative to two strands of prior research: the change-point model for NHPPs of Shen and Zhang (SeqCBS), which motivated this formulation but noted practical drawbacks of fixed binning (imprecise boundaries when breakpoints do not align with bin edges), and the continuous testing framework of Picard, Reynaud-Bouret, and Roquain (CTS), which controls FWER/FDR over a continuum of window locations via a pp-value process. The present paper instead targets the classical single-test scan statistic based on the maximum over windows, with fast calibration requiring only one sample path.

Methodology

Empirical time transformation

The key reduction uses the control sample to construct a probability integral transform. Conditional on total counts n0n_0 and n1n_1, control event times are i.i.d. from FX(t)=Λ0(t)/Λ0(1)F_X(t)=\Lambda_0(t)/\Lambda_0(1), and case times from FYF_Y. Under H0H_0 the proportionality constant λ1(t)=ρλ0(t)\lambda_1(t)=\rho\lambda_0(t)0 cancels after conditioning, so λ1(t)=ρλ0(t)\lambda_1(t)=\rho\lambda_0(t)1 and the transformed values λ1(t)=ρλ0(t)\lambda_1(t)=\rho\lambda_0(t)2 are i.i.d. Uniformλ1(t)=ρλ0(t)\lambda_1(t)=\rho\lambda_0(t)3. Since λ1(t)=ρλ0(t)\lambda_1(t)=\rho\lambda_0(t)4 is unknown, it is replaced by the empirical CDF λ1(t)=ρλ0(t)\lambda_1(t)=\rho\lambda_0(t)5 of the control sample, giving λ1(t)=ρλ0(t)\lambda_1(t)=\rho\lambda_0(t)6. The problem is thereby reduced to a scan on approximately uniform points.

Extreme value theorem

The grid-based sliding-window counts λ1(t)=ρλ0(t)\lambda_1(t)=\rho\lambda_0(t)7 form a stationary triangular array. The paper establishes three supporting results: stationarity of the count sequence (via multinomial cell-probability arguments); asymptotic distributional mixing under the conditions λ1(t)=ρλ0(t)\lambda_1(t)=\rho\lambda_0(t)8 and λ1(t)=ρλ0(t)\lambda_1(t)=\rho\lambda_0(t)9, proved via Poissonization with a Le Cam-type multinomial-to-Poisson total variation bound; and equivalence between the grid-based and continuous scan statistics with probability tending to one when ρ>0\rho>00 and ρ>0\rho>01, using a union bound on pairwise spacings near ρ>0\rho>02.

The main theorem states that, under these shrinking-window conditions together with ρ>0\rho>03,

ρ>0\rho>04

a Gumbel limit raised to the power of an extremal index ρ>0\rho>05 capturing overlap-induced clustering of exceedances. The proof combines a Cramér-type moderate-deviation approximation for binomial tails with the extreme-value theorem for stationary mixing sequences. A remark notes that the unconditional (Poissonized) version holds more directly since disjoint-interval increments are exactly independent there; the conditional fixed-ρ>0\rho>06 case is the harder one.

The assumptions yield concrete guidance: writing ρ>0\rho>07, the expected null window count should be moderately large but small relative to ρ>0\rho>08, suggesting ρ>0\rho>09. For example, with ss0, ss1 gives ss2 and ss3, consistent with the asymptotic regime, whereas ss4 gives ss5 and departs from it.

Calibration procedures

Two complementary calibrations are developed:

  • GEV calibration: GEV parameters ss6 are estimated by maximum likelihood from threshold exceedances (threshold at the 0.95 or 0.99 quantile), using a marked-Poisson-process likelihood applied to a bootstrap resample of the window-count sequence that removes dependence while preserving marginals. The extremal index is estimated by the runs estimator ss7, the ratio of exceedance clusters to exceedances. Critical values are taken as the median ss8-quantile of ss9 across bootstrap repetitions.
  • Permutation calibration: labels are randomly permuted on the pooled sample, the entire transformation and statistic are recomputed per relabeling, and the pp0-value is computed with the Phipson–Smyth correction. This provides finite-sample validity under label exchangeability without parametric assumptions.

Deletions can be handled symmetrically by scanning for unusually small counts, though the development focuses on amplifications.

Numerical results

Simulations use pp1, linear (pp2) and exponential (pp3) baselines with four pp4 combinations, window sizes pp5, 1000 Monte Carlo replications, and comparison against CTS at level pp6.

Type I error: the permutation calibration maintains empirical size close to nominal across all settings. GEV-ST shows comparable calibration but slightly inflated Type I error at the largest window size (pp7), which the authors attribute—consistently with theory—to departure from the shrinking-window regime and stronger overlap dependence. CTS is reported as more sensitive to window-size choice and baseline heterogeneity.

Power: alternatives embed a cluster by adding intensity pp8 on an interval of length equal to the scanning window. Power increases with pp9 throughout. Under linear baselines, GEV-ST and permutation consistently reach high power faster than CTS, including detection of weaker clusters. Under exponential baselines, the picture is mixed: GEV-ST and permutation dominate for small and moderate windows under slow-to-moderate growth, but for rapidly increasing baselines (n0n_00) with the largest window (n0n_01), CTS achieves higher power—one setting where the proposed methods show low power for small windows. The authors state plainly that neither calibration dominates uniformly; the choice depends on baseline shape and window size.

Application to CNV detection

The methods are applied to synthetic tumor/normal sequencing data distributed with the SeqCBS package, mimicking the HCC1954/BL1954 setting, with n0n_02 normal and n0n_03 tumor reads. Window sizes n0n_04 were chosen to satisfy the theorem's regime (n0n_05 between 0.41 and 1.65). Detected windows exceeding the critical value are merged into nonoverlapping intervals.

Across all three window sizes, GEV-ST and permutation identify the same three main CNV regions, with endpoints close to those reported by SeqCBS, and results are stable across nearby n0n_06. CTS also recovers the main regions but reports additional regions at n0n_07 and n0n_08, which the authors interpret as greater sensitivity to broader or neighboring local departures rather than as false positives per se. Sliding-window count curves show clear peaks at the SeqCBS-consistent locations for both proposed methods.

Limitations and open questions

Several limitations are acknowledged explicitly. First, the main theorem is established for the oracle transformation using the true n0n_09; the empirical-CDF replacement is validated only numerically. The authors state that validity requires the perturbation from n1n_10 not to alter the maximum window count at the order of the normalizing scale, and they leave "a full theoretical treatment of this empirical-transformation error" as future work. Second, the GEV approximation's slight Type I error inflation at larger windows means its use should be confined to settings near the shrinking-window regime, with permutation preferred when finite-sample accuracy matters. Third, power comparisons indicate no uniform winner: CTS outperforms both proposed methods under rapidly increasing exponential baselines with large windows, so the applicability boundary of the GEV/permutation approach is not fully characterized. Fourth, the extremal index is estimated with the simple runs ratio n1n_11; sensitivity to threshold choice and comparison with likelihood-based estimators (e.g., Süveges) are not examined. Finally, the analysis covers amplifications primarily, with deletions treated only by symmetry of the idea rather than evaluated empirically.

Conclusion

This paper provides a computationally efficient two-sample NHPP scan test that operates directly on point-process data without binning or parametric baseline modeling. Its theoretical contribution is a Gumbel limit with extremal-index correction for the scan statistic under a heterogeneous baseline removed by an empirical time transformation, complemented by a nonparametric permutation calibration. Simulations support reliable size control (permutation) and competitive power against continuous testing, particularly for small-to-moderate windows under slowly or moderately varying baselines, and the CNV application recovers the known regions stably across window choices. The principal unresolved issue is a rigorous account of the empirical-CDF transformation error in the extreme-value limit, along with a sharper characterization of when the asymptotic calibration loses accuracy as window size grows.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.