- The paper formulates a relationship between peak prevalence and weighted-stage aggregates in SI(k)R models, challenging the traditional factor-two heuristic.
- It rigorously analyzes both naive and Erlang scaling regimes, deriving explicit conditions and corrections that relate prevalence to easily measurable incidence data.
- The study introduces higher-order corrections that ensure less than 5% error, offering practical guidelines for epidemic forecasting in high-R0 scenarios.
Approximating Peak Prevalence in Multistage SIR Epidemics
This paper rigorously investigates the relationship between peak prevalence and weighted-stage aggregates within the SI(k)R framework, particularly under different scaling regimes for stage progression rates. The core motivation is practical: direct observation of prevalence is often infeasible due to limitations in healthcare surveillance, driving the demand for analytic proxies based on readily accessible incident case data. The authors focus on deterministic multistage SIR models, parameterized by k infectious stages, moving beyond the classical SIR (exponential infectious periods) to address realistic non-exponential dwell times.
The SI(k)R ODE system analyzed is
S˙(k)=−βS(k)I(k) I˙1(k)=βS(k)I(k)−δI1(k) I˙i(k)=δIi−1(k)−δIi(k),i=2,…,k R˙(k)=δIk(k)
with the total prevalence I(k)(t)=∑i=1kIi(k)(t) and the aggregate weighted functional
V(k)(t)=∑i=1k(k−i+1)Ii(k)(t).
The normalized version W(k)(t)=V(k)(t)/k serves as the focal analytic surrogate for prevalence, with the goal to characterize the regimes in which Wmax(k) reliably approximates Imax(k)=maxtI(k)(t).
Scaling Regimes and Analytical Results
Two distinct asymptotic regimes are studied: naive scaling (fixed per-stage rate δ as k0) and Erlang scaling (k1, so mean infectious period is fixed as k2 increases).
Naive Scaling
Under naive scaling, the infectious population concentrates at the start of the chain, yielding to leading order: k3
as k4. This directly refutes the widely used factor-two heuristic (k5) in this regime. Analytical bounds are precisely derived for k6 as functions of k7, k8, k9, and initial condition (k)0.
Erlang Scaling
By contrast, under Erlang scaling, the process admits a delay differential equation limit, with prevalence and the weighted-stage functional converging, respectively, to a moving average and a triangularly weighted moving average of incidence over the infectious period: (k)1
where (k)2 denotes incidence and (k)3 the infectious period. Under broad incidence peaks, Laplace asymptotics confirm: (k)4
and provide explicit conditions for the validity of this approximation.

Figure 1: Analytical and numerical constraints on the accuracy of the approximation (k)5.
Accuracy of the Factor-Two Approximation and Corrections
The core analytical insight is that (k)6 is valid specifically when the relative width of the incidence peak, parameterized by (k)7 (with (k)8 the local curvature of (k)9 at its maximum), is small. For sharper, high-S˙(k)=−βS(k)I(k) I˙1(k)=βS(k)I(k)−δI1(k) I˙i(k)=δIi−1(k)−δIi(k),i=2,…,k R˙(k)=δIk(k)0 outbreaks, the relative error
S˙(k)=−βS(k)I(k) I˙1(k)=βS(k)I(k)−δI1(k) I˙i(k)=δIi−1(k)−δIi(k),i=2,…,k R˙(k)=δIk(k)1
increases and explicit necessary conditions on S˙(k)=−βS(k)I(k) I˙1(k)=βS(k)I(k)−δI1(k) I˙i(k)=δIi−1(k)−δIi(k),i=2,…,k R˙(k)=δIk(k)2 for a fixed error tolerance are derived.

Figure 2: Relative errors S˙(k)=−βS(k)I(k) I˙1(k)=βS(k)I(k)−δI1(k) I˙i(k)=δIi−1(k)−δIi(k),i=2,…,k R˙(k)=δIk(k)3 of different approximations of S˙(k)=−βS(k)I(k) I˙1(k)=βS(k)I(k)−δI1(k) I˙i(k)=δIi−1(k)−δIi(k),i=2,…,k R˙(k)=δIk(k)4 as functions of S˙(k)=−βS(k)I(k) I˙1(k)=βS(k)I(k)−δI1(k) I˙i(k)=δIi−1(k)−δIi(k),i=2,…,k R˙(k)=δIk(k)5.
To extend accuracy to regimes where the factor-two rule fails, a hierarchy of higher-order corrections is constructed:
S˙(k)=−βS(k)I(k) I˙1(k)=βS(k)I(k)−δI1(k) I˙i(k)=δIi−1(k)−δIi(k),i=2,…,k R˙(k)=δIk(k)6
- Fully corrected (FC) and large-S˙(k)=−βS(k)I(k) I˙1(k)=βS(k)I(k)−δI1(k) I˙i(k)=δIi−1(k)−δIi(k),i=2,…,k R˙(k)=δIk(k)7 approximations (L(1)): explicit expressions in terms of S˙(k)=−βS(k)I(k) I˙1(k)=βS(k)I(k)−δI1(k) I˙i(k)=δIi−1(k)−δIi(k),i=2,…,k R˙(k)=δIk(k)8 only, derived from the curvature and model parameters, maintain absolute relative errors S˙(k)=−βS(k)I(k) I˙1(k)=βS(k)I(k)−δI1(k) I˙i(k)=δIi−1(k)−δIi(k),i=2,…,k R˙(k)=δIk(k)9 across broad parameter ranges.
Figure 3: Relative errors of approximations of I(k)(t)=∑i=1kIi(k)(t)0 as I(k)(t)=∑i=1kIi(k)(t)1 varies for different I(k)(t)=∑i=1kIi(k)(t)2.
Figure 4: Relative errors of parameter-based (plug-in) approximations for finite I(k)(t)=∑i=1kIi(k)(t)3 and varying I(k)(t)=∑i=1kIi(k)(t)4.
Numerical studies validate that for practical situations (epidemiologically relevant I(k)(t)=∑i=1kIi(k)(t)5), the appropriately corrected surrogate peaks (FO, FC, I(k)(t)=∑i=1kIi(k)(t)6) yield robust estimates of I(k)(t)=∑i=1kIi(k)(t)7.
Practical and Theoretical Implications
The paper has several notable implications:
- For epidemic forecasting, weighted incidence functionals, computed from easier-to-measure incidence data, can reliably estimate peak infectious burden—but only with scaling and correction formulas suited to the underlying compartmental structure.
- For mathematical epidemiology, the SII(k)(t)=∑i=1kIi(k)(t)8R limit bridges finite Markovian models and age-of-infection renewal equations, clarifying how incidence and prevalence relate as time-averaged quantities with different kernel shapes.
- The corrections and bounds derived here are essential for accurate estimation in high-I(k)(t)=∑i=1kIi(k)(t)9, sharp-peak scenarios, as encountered in rapidly spreading epidemics.
- The theoretical framework may be extended to more general dwell time distributions, heterogeneous contact networks, or partially observed processes.
Conclusion
This work provides a comprehensive asymptotic and numerical analysis of prevalence peak estimation in multistage SIR models, resolving longstanding confusion regarding the factor-two rule and developing a suite of improved surrogate approximations for practical use. The structure-function relationship between prevalence and weighted incidence, as elucidated here, supplies essential methodology for epidemic risk assessment in settings where true prevalence data are unavailable (2607.01014).