- The paper develops a two-sample nonhomogeneous Poisson process scan statistic that uses an empirical time transformation and an extremal-index-adjusted Gumbel limit to calibrate localized cluster detection without modeling the heterogeneous baseline parametrically.
- Permutation calibration maintains Type I error near the nominal 5% level across simulated baseline shapes, while GEV calibration is competitive for shrinking windows but can become slightly liberal when windows are large and overlap dependence is stronger.
- In synthetic sequencing data, both proposed methods consistently recover three main CNV regions across window sizes, matching SeqCBS closely and generally outperforming CTS for small-to-moderate windows, although CTS can be more powerful under rapidly increasing baselines with large windows.
Problem and motivation
This paper develops a scan statistic framework for detecting localized clusters in a two-sample nonhomogeneous Poisson process (NHPP) setting, motivated by copy number variation (CNV) and copy number alteration (CNA) detection in next-generation sequencing (NGS) data. The control (normal) and case (tumor) read locations are modeled as independent NHPPs on [0,1], with the null hypothesis that λ1(t)=ρλ0(t) for some ρ>0. Under the alternative, an unknown interval of length s exhibits locally elevated (amplification) or depressed (deletion) case intensity. The central inferential difficulty is calibrating the tail of the scan statistic—the maximum count over sliding windows—when the baseline intensity is heterogeneous, unknown, and induces strong dependence among overlapping windows. Existing rigorous null distributions for scan statistics under NHPP models are limited, particularly in the two-sample setting; homogeneous-baseline approximations inflate Type I error because large local counts may reflect baseline heterogeneity rather than a true cluster.
The work is positioned relative to two strands of prior research: the change-point model for NHPPs of Shen and Zhang (SeqCBS), which motivated this formulation but noted practical drawbacks of fixed binning (imprecise boundaries when breakpoints do not align with bin edges), and the continuous testing framework of Picard, Reynaud-Bouret, and Roquain (CTS), which controls FWER/FDR over a continuum of window locations via a p-value process. The present paper instead targets the classical single-test scan statistic based on the maximum over windows, with fast calibration requiring only one sample path.
Methodology
The key reduction uses the control sample to construct a probability integral transform. Conditional on total counts n0 and n1, control event times are i.i.d. from FX(t)=Λ0(t)/Λ0(1), and case times from FY. Under H0 the proportionality constant λ1(t)=ρλ0(t)0 cancels after conditioning, so λ1(t)=ρλ0(t)1 and the transformed values λ1(t)=ρλ0(t)2 are i.i.d. Uniformλ1(t)=ρλ0(t)3. Since λ1(t)=ρλ0(t)4 is unknown, it is replaced by the empirical CDF λ1(t)=ρλ0(t)5 of the control sample, giving λ1(t)=ρλ0(t)6. The problem is thereby reduced to a scan on approximately uniform points.
Extreme value theorem
The grid-based sliding-window counts λ1(t)=ρλ0(t)7 form a stationary triangular array. The paper establishes three supporting results: stationarity of the count sequence (via multinomial cell-probability arguments); asymptotic distributional mixing under the conditions λ1(t)=ρλ0(t)8 and λ1(t)=ρλ0(t)9, proved via Poissonization with a Le Cam-type multinomial-to-Poisson total variation bound; and equivalence between the grid-based and continuous scan statistics with probability tending to one when ρ>00 and ρ>01, using a union bound on pairwise spacings near ρ>02.
The main theorem states that, under these shrinking-window conditions together with ρ>03,
ρ>04
a Gumbel limit raised to the power of an extremal index ρ>05 capturing overlap-induced clustering of exceedances. The proof combines a Cramér-type moderate-deviation approximation for binomial tails with the extreme-value theorem for stationary mixing sequences. A remark notes that the unconditional (Poissonized) version holds more directly since disjoint-interval increments are exactly independent there; the conditional fixed-ρ>06 case is the harder one.
The assumptions yield concrete guidance: writing ρ>07, the expected null window count should be moderately large but small relative to ρ>08, suggesting ρ>09. For example, with s0, s1 gives s2 and s3, consistent with the asymptotic regime, whereas s4 gives s5 and departs from it.
Calibration procedures
Two complementary calibrations are developed:
- GEV calibration: GEV parameters s6 are estimated by maximum likelihood from threshold exceedances (threshold at the 0.95 or 0.99 quantile), using a marked-Poisson-process likelihood applied to a bootstrap resample of the window-count sequence that removes dependence while preserving marginals. The extremal index is estimated by the runs estimator s7, the ratio of exceedance clusters to exceedances. Critical values are taken as the median s8-quantile of s9 across bootstrap repetitions.
- Permutation calibration: labels are randomly permuted on the pooled sample, the entire transformation and statistic are recomputed per relabeling, and the p0-value is computed with the Phipson–Smyth correction. This provides finite-sample validity under label exchangeability without parametric assumptions.
Deletions can be handled symmetrically by scanning for unusually small counts, though the development focuses on amplifications.
Numerical results
Simulations use p1, linear (p2) and exponential (p3) baselines with four p4 combinations, window sizes p5, 1000 Monte Carlo replications, and comparison against CTS at level p6.
Type I error: the permutation calibration maintains empirical size close to nominal across all settings. GEV-ST shows comparable calibration but slightly inflated Type I error at the largest window size (p7), which the authors attribute—consistently with theory—to departure from the shrinking-window regime and stronger overlap dependence. CTS is reported as more sensitive to window-size choice and baseline heterogeneity.
Power: alternatives embed a cluster by adding intensity p8 on an interval of length equal to the scanning window. Power increases with p9 throughout. Under linear baselines, GEV-ST and permutation consistently reach high power faster than CTS, including detection of weaker clusters. Under exponential baselines, the picture is mixed: GEV-ST and permutation dominate for small and moderate windows under slow-to-moderate growth, but for rapidly increasing baselines (n00) with the largest window (n01), CTS achieves higher power—one setting where the proposed methods show low power for small windows. The authors state plainly that neither calibration dominates uniformly; the choice depends on baseline shape and window size.
Application to CNV detection
The methods are applied to synthetic tumor/normal sequencing data distributed with the SeqCBS package, mimicking the HCC1954/BL1954 setting, with n02 normal and n03 tumor reads. Window sizes n04 were chosen to satisfy the theorem's regime (n05 between 0.41 and 1.65). Detected windows exceeding the critical value are merged into nonoverlapping intervals.
Across all three window sizes, GEV-ST and permutation identify the same three main CNV regions, with endpoints close to those reported by SeqCBS, and results are stable across nearby n06. CTS also recovers the main regions but reports additional regions at n07 and n08, which the authors interpret as greater sensitivity to broader or neighboring local departures rather than as false positives per se. Sliding-window count curves show clear peaks at the SeqCBS-consistent locations for both proposed methods.
Limitations and open questions
Several limitations are acknowledged explicitly. First, the main theorem is established for the oracle transformation using the true n09; the empirical-CDF replacement is validated only numerically. The authors state that validity requires the perturbation from n10 not to alter the maximum window count at the order of the normalizing scale, and they leave "a full theoretical treatment of this empirical-transformation error" as future work. Second, the GEV approximation's slight Type I error inflation at larger windows means its use should be confined to settings near the shrinking-window regime, with permutation preferred when finite-sample accuracy matters. Third, power comparisons indicate no uniform winner: CTS outperforms both proposed methods under rapidly increasing exponential baselines with large windows, so the applicability boundary of the GEV/permutation approach is not fully characterized. Fourth, the extremal index is estimated with the simple runs ratio n11; sensitivity to threshold choice and comparison with likelihood-based estimators (e.g., Süveges) are not examined. Finally, the analysis covers amplifications primarily, with deletions treated only by symmetry of the idea rather than evaluated empirically.
Conclusion
This paper provides a computationally efficient two-sample NHPP scan test that operates directly on point-process data without binning or parametric baseline modeling. Its theoretical contribution is a Gumbel limit with extremal-index correction for the scan statistic under a heterogeneous baseline removed by an empirical time transformation, complemented by a nonparametric permutation calibration. Simulations support reliable size control (permutation) and competitive power against continuous testing, particularly for small-to-moderate windows under slowly or moderately varying baselines, and the CNV application recovers the known regions stably across window choices. The principal unresolved issue is a rigorous account of the empirical-CDF transformation error in the extreme-value limit, along with a sharper characterization of when the asymptotic calibration loses accuracy as window size grows.