Determine whether lead-time-specific architecture advantages are masked by averaged metrics

Determine whether the five significant-wave-height forecasting architectures—DLinear, LSTM, PatchTST, ResAttLstm, and Mamba2—exhibit performance differences at particular 1–6-hour lead times that are obscured when metrics are averaged across all six lead times.

Background

All reported metrics in the study are averaged over the six forecast lead times from 1 to 6 hours. Such aggregation may conceal different behavior at individual horizons, including a possible advantage for recurrent models at short lead times or attention-based models at longer horizons.

The paper explicitly leaves this issue unresolved and identifies lead-time-disaggregated analysis as necessary to establish whether architecture choice matters differently across forecast horizons.

References

It remains possible that architectural differences manifest at specific lead times---for instance, the relative advantage of recurrent architectures at short lead times versus attention-based models at longer horizons---but are obscured by averaging; lead-time-disaggregated analysis is left to future work.

— On the Limits of Univariate Deep Learning for Significant Wave Height Forecasting  (2609.30688 - Zhai et al., 25 Sep 2026) in Section 4.3, “Limitations and generalisability”