Papers
Topics
Authors
Recent
Search
2000 character limit reached

Horizon-Aware Forecasting of Passenger Assistance Demand for Rail Station Workforce Planning

Published 8 Apr 2026 in stat.AP and cs.LG | (2604.16464v1)

Abstract: Passenger assistance services are essential for accessible rail travel, yet demand varies substantially across stations and over time, creating challenges for workforce planning and staff rostering. This paper presents a data-driven decision support framework for forecasting station-level passenger assistance demand and translating forecasts into workforce plans. The forecasting component applies a horizon-aware Prophet modelling approach using multi-source operational data, while the planning component maps demand forecasts to staffing requirements under service and operational constraints through an interpretable red-amber-green risk framework. The approach has been implemented within a production-grade system to support routine planning and staffing decisions across LNER-managed stations. Results demonstrate improved forecast accuracy relative to year-on-year baseline methods, with absolute error reduced by up to 76.9%, and show that forecast-informed staffing is associated with an approximate 50% reduction in failed passenger assistance deliveries attributable to staff availability. These findings highlight the value of integrating interpretable forecasting with operational work.

Summary

  • The paper introduces horizon-bucket Prophet models with leakage-free booking features to forecast hourly departure and arrival assistance demand at different planning lead times.
  • Forecast accuracy improved consistently as the horizon shortened, with error reductions of up to 76.9% and coverage rising to 96–100% at representative stations.
  • The deployed system translates forecasts into transparent red–amber–green staffing alerts and was associated with roughly 50% fewer staff-availability-related assistance failures, although evidence remains preliminary.

Operational context and motivation

Passenger assistance services are a regulatory obligation for UK train operating companies, yet national reporting indicates that only around 78% of passengers receive all booked assistance, and delivery failures carry financial, reputational, and accessibility costs. This paper, authored by researchers at London North Eastern Railway (LNER), addresses a gap between the extensive literature on aggregate passenger demand forecasting and the near-absence of quantitative forecasting work on station-level assistance demand — a quantity that is not a fixed proportion of passenger flows and varies substantially across stations. The authors present an end-to-end decision-support system deployed across LNER's thirteen managed stations, combining horizon-aware Prophet forecasting of hourly assistance demand with an interpretable workforce capacity assessment.

The operational setting motivates the design choices directly. Booking lead-time analysis shows approximately 44% of assistance bookings are made within the final week before travel, while rosters are fixed roughly four weeks ahead and later adjustments rely on increasingly expensive reactive measures (overtime, rest-day working). Forecast value therefore depends on when information is available relative to each decision point, which the authors treat as a first-class modelling constraint rather than an afterthought.

Horizon-bucket forecasting methodology

The core methodological contribution is a horizon-bucket routing scheme over decomposable additive (Prophet-style) models. The forecast target is the hourly count of pre-booked assistance events per station, with journey-level requests decomposed into departure (DEP) and arrival (ARR) events at each involved station. Five non-overlapping horizon buckets — very short-term (1–2 days), short-term (3–7), medium I (8–14), medium II (15–28), and long-term (>28) — each receive a separately trained model, with future timestamps deterministically routed by their lead time H=tt0H = t - t_0.

Each bucket-specific model follows the standard decomposition:

y(t)=g(t)+jsj(t)+h(t)+xk(t)βk+ε(t)y(t) = g(t) + \sum_{j} s_j(t) + h(t) + \mathbf{x}_k(t)^\top \boldsymbol{\beta}_k + \varepsilon(t)

with piecewise-linear trend, daily/weekly/yearly seasonalities, holiday effects, and bucket-specific regressors. Two design elements are notable:

  • As-of booking features. Cumulative booking counts at multiple lead-time thresholds τ\tau, plus adjacent-threshold differences capturing late-booking pace, are constructed strictly relative to the forecast origin. This prevents leakage and ensures reported performance reflects what planners would actually know at each decision point.
  • Domain-specific holiday windows. Rather than treating Christmas as a single-day effect, the model uses lead/lag windows around the year-end period, Easter, and other high-impact dates, accommodating year-to-year shifts in peak timing driven by the day-of-week on which holidays fall.

TUAG records are deliberately excluded from the training target because disruption-driven reclassification introduces labelling artefacts; instead, a historical seasonal TUAG rate is supplied to managers as a separate indicative estimate. Weather covariates (temperature, rainfall, humidity) are joined at hourly resolution. Evaluation uses MAE, an asymmetric RMSE penalising under-prediction at twice the weight of over-prediction (reflecting the higher operational cost of understaffing), and a station-specific coverage probability within tolerance bands.

From forecasts to staffing assessment

The planning component is explicitly framed as decision support rather than optimisation. Hourly effective capacity is derived from the base roster, separating primary staff (Performance Station Assistants, Station Support Customer Service) from secondary roles (dispatch-focused SCSA, help-point SSA) whose availability is discounted by role-specific factors αi\alpha_i. Capacity is computed as:

Ctotal,h=NP,hAh(1M)+iNi,hαiAh(1M)C_{\text{total},h} = N_{P,h}\, A_h(1-M) + \sum_i N_{i,h}\,\alpha_i\, A_h(1-M)

with parameters set from operational expertise: Ah=4A_h = 4 assists/hour, margin M=0.10M = 0.10, and secondary availability α=0.30\alpha = 0.30. Each station-hour is then classified via a red–amber–green rule: green if forecast demand fits within primary capacity, amber if secondary capacity is required, red if total capacity is exceeded. Roles usable only as a last resort (Information Controllers, Duty Team Leaders) are excluded so that estimated capacity reflects routinely deployable resource.

Forecast accuracy results

Evaluation covers three representative stations spanning contrasting demand regimes: London King's Cross (high-volume long-distance terminal), York (interchange hub), and Berwick-upon-Tweed (low-volume, sparse demand). Across all stations and both hourly and daily resolutions, error decreases monotonically as the horizon shortens. At the daily level, very short-term forecasts reduce aRMSE relative to the year-on-year baseline by approximately 66% at King's Cross, 74% at York, and 66% at Berwick-upon-Tweed; the abstract reports absolute error reductions of up to 76.9%. Representative figures:

Station Baseline MAE Very short-term MAE Baseline coverage Very short-term coverage
King's Cross (daily) 37.1 8.6 58.4% 96.3%
York (daily) 24.2 6.9 72.0% 99.7%
Berwick-upon-Tweed (daily) 4.8 1.4 100% 100%

Coverage improvements are operationally significant at King's Cross, where the baseline misses tolerance in over 40% of periods but the very short-term model achieves 97.9% coverage at hourly resolution. The consistency of monotonic improvement across sparse-demand Berwick-upon-Tweed suggests the approach generalises beyond high-volume contexts rather than being driven by a single demand profile.

Explainability and residual diagnostics

The GAM-style decomposition provides component-level attribution: trend, weekly/yearly/intraday seasonalities, holiday windows, and regressor contributions can each be inspected, allowing planners to distinguish structural seasonality from emerging booking activity. Residual diagnostics for York reveal systematic deviations coinciding with major engineering works not represented in the model specification. The authors treat these as actionable diagnostics identifying candidate exogenous regressors (engineering works, disruption indicators) rather than unexplained model failure — a candid acknowledgement that the current specification omits known demand drivers.

Post-deployment impact

The system entered production in December 2025. Early evidence associates its use with an approximate 50% reduction in failed assistance deliveries attributable to staff availability, measured per 1,000 assists against the equivalent pre-deployment period. A concrete validation case arose in the post-Christmas period: when 27 December fell on a Saturday, demand redistributed across 27–29 December rather than peaking on a single day. The year-on-year baseline would have underestimated 29 December demand by more than 700 bookings network-wide, whereas the deployed framework reduced forecast error by approximately 93% relative to the static baseline, enabling proactive roster adjustments. The authors appropriately caveat these as early results from a limited post-deployment window rather than a long-term evaluation, and the operational association is observational rather than experimentally controlled.

Limitations and open questions

Several limitations are conceded within the paper. Staffing parameters (AhA_h, MM, y(t)=g(t)+jsj(t)+h(t)+xk(t)βk+ε(t)y(t) = g(t) + \sum_{j} s_j(t) + h(t) + \mathbf{x}_k(t)^\top \boldsymbol{\beta}_k + \varepsilon(t)0) rest on operational expertise without calibration or sensitivity analysis, so the RAG classifications inherit whatever bias these assumptions carry. Engineering works, large-scale events, and real-time disruption indicators are absent from the current regressor set despite demonstrably driving residual structure. The post-deployment service-reliability improvement lacks a controlled counterfactual, and evaluation is restricted to three representative stations out of thirteen, leaving open whether reported gains hold uniformly across the full portfolio. Finally, the framework models only pre-booked demand; TUAG volumes enter only through a static seasonal rate, so unbooked demand surges under disruption remain outside the forecast's scope.

Conclusion

This paper demonstrates a production-deployed integration of interpretable, horizon-aware time-series forecasting with operational workforce planning for rail passenger assistance. Its principal contributions are the horizon-bucket routing aligned to staggered planning responsibilities, leakage-free as-of feature construction, and a transparent RAG translation of forecasts into staffing risk. Reported results — up to 76.9% error reduction versus year-on-year baselines, ~93% error reduction in a calendar-shift stress case, and an approximate halving of staff-availability-related delivery failures after deployment — indicate tangible operational value, though the evidence base remains early-stage and several model inputs await empirical calibration.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.