Effect of Context Length on SimpleTimeBench Performance

Determine how varying context lengths affect the performance of time series foundation models on the SimpleTimeBench forecasting tasks, including the minimum historical window required to observe the relevant patterns.

Background

SimpleTimeBench evaluates time series foundation models on controlled synthetic processes whose underlying forecasting behavior is known, including deterministic trends, periodic signals, noise processes, and stochastic correlated processes. The reported experiments use a fixed context length of 208 time steps, together with a forecast horizon of 48 steps and datasets of length 256 time steps.

The paper notes that an initial investigation observed similar model characteristics across a range of parameter settings, but the systematic effect of context length was not studied. Establishing how performance changes with the available historical window would clarify the robustness of the benchmark results and identify the minimum context needed for models to recognize and forecast the tested temporal primitives.

References

TSFMs should be able to solve these tasks at a wide variety of context lengths, up to a minimum where there is insufficient historical window to observe; we leave such an investigation of the effects of context length for future work.

— Foundations without Fundamentals: Zero-Shot Blind Spots in Time Series FMs  (2610.02058 - Ghoroghchian et al., 1 Oct 2026) in Appendix, Section \ref{apx:simpletime}, subsection “Results”