- The paper develops CAIRN, a strictly causal meteorology-driven framework that combines transfer entropy with a physics-anchored S4D state-space model to nowcast fugitive landfill H₂S emissions at 15-minute resolution.
- The framework identifies wind direction, wind speed, and pressure as core causal drivers, while the S4D nowcaster outperforms XGBoost for high-risk detection with F₁ 0.533, recall 0.618, and AUC 0.898.
- The multi-site alert classifier achieves weighted κ 0.709 against concentration-based tiers and correlates with held-out daily community odour complaints at r = 0.729, supporting proactive public-health responses while requiring site-specific validation.
Overview and study context
This paper develops and validates a meteorology-driven nowcasting framework ("CAIRN") for fugitive landfill emissions of hydrogen sulphide (H₂S), motivated by a European municipal landfill subject to a multi-year regulatory and public-health response. The site accepted predominantly non-hazardous waste with one gypsum-containing cell; from 2021 onward community complaints rose substantially, 24-hour mean H₂S approached but generally remained below the ATSDR Intermediate Minimal Risk Level (30μgm−3), while concentrations frequently exceeded the WHO 30-minute odour-annoyance guideline value of 7μgm−3. The regulator closed the site at the end of 2024. The paper's central claim is that a strictly causal, meteorology-only model can anticipate receptor-level exposure at 15-minute resolution well enough to support proactive public-health action under the FIDOL framework and UK statutory-nuisance structures, rather than the retrospective source–pathway–receptor assessment that currently dominates practice.
The work is organised around three methodological pillars: an extended phenomenological characterisation of the emission record; a multiscale transfer entropy (MSTE) causal analysis whose output is used as an architectural prior; and a dual-pathway S4D state-space nowcaster benchmarked against engineered XGBoost classifiers, culminating in a Bayesian multi-site alert-tier classifier validated against held-out community odour-complaint data.
Environmental characterisation of the emission record
Using n=35,041 valid 15-minute H₂S observations at the principal receptor (MMF9) across 2024, the characterisation establishes source attribution, chemical fingerprinting and meteorological mechanism. Conditional probability function analysis shows exceedance probability peaking in the west-northwest sector at bearing 305∘ (CPF ≈0.35), with a peak/opposing-sector ratio of 4.09×, and CPF cones from three stations converging on a common source area — spatial triangulation of the emitting region. The H₂S–CH₄ co-emission fingerprint is strong: Pearson r=0.832, Jaccard spike co-occurrence J=0.654 (1,386 co-occurring spikes versus 87.7 expected under independence), and a super-linear power-law exponent b=1.96, which the authors interpret as consistent with sulphate-reducer activity in deeper anaerobic cells, while conceding that the downwind exponent does not map one-to-one onto subsurface microbial kinetics. Negative controls are clean: PM₁₀ and PM₂.₅ show no positive association, and the moderate NOₓ correlation (r=0.377) is attributed to joint nocturnal trapping given a 6-hour diurnal phase offset.
The diurnal and seasonal structure is pronounced: mean concentrations peak at 02:00 LST (7μgm−30) against a mid-afternoon trough of 7μgm−31 (a 7μgm−32 ratio), February exceeds August by 7μgm−33, and sunrise-aligned aggregation shows a 7μgm−34 pre-/post-sunrise reduction reproduced across all 365 days — evidence that inversion break-up, not episodic operations, governs the cycle. The temperature–H₂S phase loop shows strong counter-clockwise hysteresis (7μgm−35), whereas the wind-speed loop shows none, indicating wind dilution acts without memory at diurnal timescales. Individual linear drivers explain little variance (7μgm−36–7μgm−37), but extreme spikes cluster tightly in the low-wind, low-temperature, negative-tendency region corresponding to Pasquill E–F stable boundary layers. This motivates the multivariate, nonlinear modelling that follows.
Causal hierarchy via multiscale transfer entropy
The MSTE analysis applies conditional transfer entropy (Frenzel–Pompe KSG estimator with Theiler window 7μgm−38, rank transform, adaptive 7μgm−39, single-confounder conditioning) across temporal scales from 15 min to 12 h, with significance from n=35,0410 circular-shift surrogates and Benjamini–Hochberg correction. Three drivers form the causal core: wind direction, wind speed and atmospheric pressure, significant at all scales n=35,0411 h. Pressure derivatives become causally significant only at n=35,0412 h, which the authors identify as the information-theoretic counterpart of barometric-pumping experiments showing 20-fold gas breakthroughs over sub-daily timescales. Temperature is significant only at the 6 h scale, identifying the site as a nocturnal-accumulation regime where solar heating matters as integrated forcing rather than an instantaneous driver.
Two robustness analyses deserve emphasis. A six-configuration estimator ablation shows the core-driver result survives every methodological choice, that bivariate (unconditioned) transfer entropy inflates effect sizes and false positives, and that history embedding depth is the dominant sensitivity axis. A split-half resampling (January–June vs July–December) reproduces the ordering WD > pressure > wind speed identically in both halves, with sign agreement in 19 of 21 cells. However, the paper is candid about two weaknesses: coarse-scale cells (n=35,0413 h) rest on partially reconstructed blocks because no block is gap-free, and the full-year estimates exceed both half-year estimates for pressure and wind speed — behaviour consistent with sample-size-dependent estimator bias, so magnitudes should be read as ordinal rather than absolute. Under a stricter Benjamini–Yekutieli correction, 20 of 29 significant cells survive, including all wind-direction and pressure cells.
The architectural contribution is a dual-pathway diagonal S4 model whose kernel timescales are initialised on the causally significant horizons identified by MSTE: fast lanes anchored at n=35,0414 h and slow lanes at n=35,0415 h, the longest scale at which pressure remains causally significant. Training uses focal cross-entropy with High-class weighting n=35,0416, AdamW with five differential learning-rate groups protecting the timescale prior, warmup-freeze of the n=35,0417 spectrum, and strict causality enforced at kernel, convolution and pooling levels. Evaluation is a 13-week walk-forward protocol over October–December 2024 (n=35,0418 timesteps) with five explicit no-leakage guarantees.
A six-condition initialisation ablation isolates the value of the physics anchor: the MSTE-aligned configuration achieves period-wide F₁-High of n=35,0419 versus 305∘0–305∘1 for alternatives, with paired per-fold Wilcoxon tests reaching Benjamini–Hochberg-corrected 305∘2 (weakest against random initialisation, which does not reach significance on a sign test). Against the matched XGBoost benchmark, the S4 model leads on every High-class detection metric — F₁-High 305∘3 vs 305∘4, recall 305∘5 vs 305∘6, AUC 305∘7 vs 305∘8 — contributing 38 additional true positives with 14 fewer false positives. The individual margins are modest and do not reach significance on paired per-fold tests (F₁-High 305∘9); only the pooled McNemar test over all predictions is significant (≈0.350). Notably, probabilistic ranking metrics favour the benchmark once checkpoint selection is removed, reflecting a better-ranked but less well-calibrated posterior for the S4 model. A three-arm decomposition attributes the gains to memory: hand-coded engineering adds ≈0.351 F₁ over a memoryless floor and learnt state-space memory adds ≈0.352, both clearing the paired Wilcoxon test, though neither clears a sign test. The slow-lane kernel also reproduces marginal response curves for ≈0.353 and ≈0.354 without those derivatives appearing as inputs, although the lane-resolved association is weak (≈0.355) and the lane-attribution itself reverses depending on whether the ablation mask is applied at the kernel or block level — reported as consistency, not mechanistic attribution.
Multi-site Bayesian alert-tier classifier and community validation
Four independently trained walk-forward nowcasters (H₂S and CH₄ at two receptors) feed a Bayesian log-odds accumulator with asymmetric hysteresis latching (45-min onset, 30-min clearance matched to WHO averaging), producing a continuous network-event posterior mapped to four ordinal tiers aligned with the regulatory decision landscape. The predicted tier is strictly meteorology-only; raw concentrations enter only the deterministic ground truth. On the synchronous January–March 2025 window (≈0.356 timesteps), the classifier achieves quadratic-weighted Cohen's ≈0.357 (95% CI 0.557–0.852) against the ground-truth tier, with Tier 0 F₁ of ≈0.358 and Tier 3 F₁ of ≈0.359 (recall 0.493). The improvement over a same-architecture vote-count baseline is 4.09×0 (paired block bootstrap CI 0.025–0.109, 4.09×1), though the predicted tier coincides with vote count minus one on 94.2% of timesteps — a calibrated refinement rather than a categorically different rule.
The most consequential result is external validation against daily community odour complaints (4.09×2 days), a record fully held out from training, hysteresis selection and likelihood-ratio design. Daily-mean predicted tier correlates with complaint counts at Pearson 4.09×3 (CI 0.549–0.837, 4.09×4), capturing approximately 84% of the ground-truth ceiling's explained variance, with a unique lag-0 cross-correlation peak confirming nowcast alignment. This demonstrates that machine-learning probability outputs can be made directly interpretable against the FIDOL framework underpinning UK statutory-nuisance assessment — the operational target being the WHO annoyance guideline, orders of magnitude below acute-exposure thresholds.
Limitations and open questions
The paper concedes several material constraints. The headline 4.09×5 is a mixed quantity: fusion parameters were tuned on nine of thirteen weeks, and a stratified held-out subset yields 4.09×6; leave-one-week-out refitting bounds tuning optimism at only 0.006, but the wide block-bootstrap interval means operational planning should use the lower bound. Tier cut-points are operational choices not reproduced by the documented grid search, though substitution of the refit values changes 4.09×7 by at most 0.013. The conditional-independence assumption underlying the naive Bayes fusion is violated by shared chemistry, overlapping advection fields and common local meteorology; the deliberately conservative likelihood ratios absorb this correlation, and a copula or empirical-Bayes treatment is left open. Per-tier performance at the extremes rests on roughly 30–100 effective independent episodes, and intermediate tiers carry most off-by-one disagreement (Tier 1 precision 0.266). The causal hierarchy is explicitly regime-specific and must be re-estimated per site given documented inter-site variability in H₂S/CH₄ ratios spanning three orders of magnitude. Finally, conversion to genuine forecasting requires substituting numerical-weather-prediction meteorology at inference — architecturally trivial given strict causality, but the operational pipeline is unbuilt, and coupling to clinical endpoints beyond complaint counts remains untested.
Conclusion
The paper presents a coherent pipeline from phenomenological characterisation through information-theoretic causal analysis to a physics-anchored state-space nowcaster and a calibrated multi-site alert-tier classifier, evaluated under leak-free walk-forward protocols with unusually explicit accounting of tuning optimism, estimator degeneracy and selection-protocol effects. Its strongest quantitative claims are the external complaint correlation (4.09×8 at lag 0) and the monotone benefit of learnt memory over both memoryless and hand-engineered baselines; its weakest are the modest, individually non-significant margins of the S4 architecture over a well-tuned XGBoost benchmark and the dependence of headline agreement figures on partially in-sample fusion parameters. The demonstrated feasibility of meteorology-only, community-endpoint-calibrated nowcasting establishes a concrete basis for anticipatory public-health response at fugitive-emission sites, contingent on auditable monitoring infrastructure and site-specific re-estimation of the causal hierarchy.