---
title: Multi-Horizon Surrogates
url: https://www.emergentmind.com/topics/multi-horizon-surrogates
type: topic
---

# Multi-Horizon Surrogates

A multi-horizon surrogate is a learned model designed to approximate states, outputs, or objective values across a range of temporal or scenario-based prediction horizons within a single architecture. Such surrogates address the limitations of single-horizon models, providing stability, adaptivity, and efficiency for tasks that require forecasting, optimization, or causal inference over multiple, potentially variable forward steps. Notable applications include time-marching turbulent flow prediction, stochastic programming under uncertainty, physical dynamics modeling, and estimation of long-term causal effects from short-term proxies.

## 1. Formal Definitions and Conceptual Scope

A multi-horizon surrogate seeks to learn, within a single model or closely coupled ensemble, a family of parameterized operators $\{\widehat{\mathcal{T}_\Delta}\}$ or functions $f_\theta(\cdot, h)$ such that, for any valid horizon $h$ (time step $\Delta$, integer stride, or generalized decision horizon), the output approximates the system’s evolution at that horizon, i.e., $\mathbf{x}(t+h) \approx \widehat{\mathcal{T}_h}[\mathbf{x}(t)]$ or for general physical/abstract state $s$, $\hat s_{t+T} = f_\theta(s_t, T) \approx \Phi_T(s_t)$. This contrasts with classical surrogates trained for a fixed interval, which may accumulate substantial error or fail to generalize if chained or interpolated to new horizons.

Multi-horizon surrogates have been proposed and evaluated for:

- Autoregressive and direct prediction of physical states at variable temporal strides [2604.12794], [2605.28317].
- Surrogate modeling of recourse objectives in stochastic programs, embedding the surrogate as a functional approximation of expected cost over scenario trees [2512.02294].
- Estimation of long-term causal effects using surrogates constructed from short-term outcomes and their latent or proxy variables [2208.04589].

## 2. Architectures and Methodological Principles

A core requirement is conditioning the surrogate on the desired prediction or optimization horizon. Several conditioning and architectural motifs are prominent:

**Conditioning Mechanisms**

- **Stride/Time-Step Routing**: Explicit input of the stride $\Delta$ or horizon $T$ into the model, followed by embedding (e.g., via small fully connected networks), enables the network to generalize over discrete or continuous horizon values [2604.12794], [2605.28317].
- **Mixture-of-Experts (MoE) Structures**: A multi-step-size mixture-of-experts neural operator uses dyadic (power-of-two) stride-specific experts $E_k$, a shared expert $E_0$, and soft routing weights $r_k(\Delta)$ derived from $\Delta$ (typically in log-space with soft-kernel Gaussian blending) to interpolate and fuse predictions across stride scales [2604.12794].
- **Feature-wise Linear Modulation (FiLM) Layers**: Embedding the horizon $T$ and modulating internal representations at every block through FiLM scaling/shifting allows for horizon-aware adaptation without an explosion in parameters [2605.28317].

**Architectural Backbones and Training**

- **Implicit Factorized Transformers (IFactFormer-m)**: Used for spatiotemporal map prediction in turbulence, employing axial factorized attention (per axis), input lifting, and deep-equilibrium residual iterations for stability in long rollouts [2604.12794].
- **U-Net and Residual MLP**: In world model surrogates for PDE and ODE domains, standard convolutional and residual multilayer architectures are used, with horizon conditioning applied at every stage [2605.28317].
- **Feed-forward Neural Networks**: In stochastic programming, ReLU-activated FFNNs are trained to map scenario-level first-stage decisions to operational recourse objective predictions, embedded directly into the mixed-integer linear program [2512.02294].
- **Identifiable VAEs (iVAE)**: For causal effect estimation using multi-horizon surrogates, variational autoencoders recover latent surrogates from mixed observed/proxy short-term outcomes, supporting recoverability and unbiased estimation [2208.04589].

## 3. Training and Dataset Construction

Robust multi-horizon generalization hinges on comprehensive exposure during training:

- **Multi-Stride/Ladder Sampling**: Training tuples $(\mathbf{u}^n, \mathbf{u}^{n+s}, s)$ or $(s_0, s_T, T)$ are sampled with $s$ or $T$ drawn uniformly over a geometric set of strides (e.g., $\{1,2,\ldots,T_{\max}\}$ in turbulence; $\{1,2,4,8,\ldots,64\}$ in physical world models), enforcing coverage of both fine and coarse prediction intervals [2604.12794], [2605.28317].
- **Direct Supervision and DAgger Refinement**: Models are supervised on reference solver output at multiple horizons; DAgger-style rollin policies introduce on-policy samples to remedy covariate shift and bottleneck error propagation [2605.28317].
- **Surrogate Cost Data**: In stochastic programs, training data are gathered by Latin hypercube sampling of feasible first-stage decisions and solving the full recourse problem to record target costs, enabling accurate regression of the neural surrogate objective surface [2512.02294].
- **Multi-set Fusion for Causal Surrogacy**: Joint use of observational and experimental data, encompassing both observed surrogates and proxy measurements, enables identification and representation of the full surrogate variable mediating long-term effects [2208.04589].

## 4. Applications and Empirical Performance

Multi-horizon surrogates have demonstrated performance advantages across a breadth of domains:
- **Turbulent Flow Simulation**: The Ms-MoE-IFactFormer operator yields stable, long-horizon rollouts over tens of thousands of fine time steps, with 30–50% lower late-time $L^2$ error growth and preservation of energy spectra with $<5\%$ bias at high wavenumbers [2604.12794].
- **Physical World Dynamics**: Horizon-conditioned surrogates predict future physical state in one forward pass, achieving CPU speedups of 26–72$\times$ relative to PDE solvers, and offering trustworthy error detection via step-doubling error maps (AUROC up to 0.98 on challenging shock regions) [2605.28317].
- **Stochastic Programming**: Embedded neural network surrogates yield up to 34.7$\times$ faster solve times in multi-horizon energy planning, with out-of-sample cost generalization surpassing deterministic equivalents, and mean absolute percentage error under 2.5% even for large scenario banks [2512.02294].
- **Causal Inference**: The LASER iVAE-based estimator attains uniformly lower MAPE on average treatment effect estimation than classical and deep-learning baselines, demonstrating the utility of latent multi-horizon surrogates in unbiased long-term causal inference, especially when observed proxies are noisy [2208.04589].

## 5. Limitations and Theoretical Considerations

Key limitations and open questions highlighted across studies include:
- **Training Cost and Scalability**: Multi-horizon surrogates, especially those using mixture-of-experts or embedded neural architectures, require increased training time and GPU memory; however, the resulting models are flexible and reusable, partially amortizing these costs over repeated deployment [2604.12794], [2512.02294].
- **Dyadic Partition Effects**: The use of a fixed dyadic (power-of-two) horizon partition in MoE architectures may result in weaker specialization at boundaries; possible remedies include learned or nonuniform binning [2604.12794].
- **Compositional Consistency**: Surrogates may accrue error over successive compositions; incorporating consistency losses or semigroup constraints ($\widehat{G}(\widehat{G}(u,s_1),s_2) \approx \widehat{G}(u,s_1+s_2)$) could mitigate long-horizon drift [2604.12794], [2605.28317].
- **Proxy and Surrogate Identification**: In settings with mixed observed and latent surrogates, identifiability is only guaranteed under specific exponential-family assumptions and sufficient experiment–observational overlap; applications in high-dimensional or partially aligned datasets may challenge these assumptions [2208.04589].

## 6. Impact, Generality, and Future Directions

Multi-horizon surrogates have shifted the paradigm in simulation, optimization, and inference from fixed-step, single-output models toward unified architectures that faithfully extrapolate across temporal or scenario horizons. This capability facilitates:
- Efficient uncertainty quantification and robust planning in large-scale stochastic programs without enumerating the full recourse structure [2512.02294].
- Extensible modeling of physical systems with variable-step integration and adaptive solver fallback, including formal error quantification even at discontinuities [2605.28317].
- Causal effect estimation in domains where only short-term or partial surrogates are available, leveraging latent disentanglement for unbiased inference [2208.04589].
- Stable simulation of high-dimensional, chaotic systems (e.g., 3D turbulence) for far longer horizons and at finer resolutions than previous neural operator surrogates [2604.12794].

Future research directions include learned time/space partitioning, compositional invariance penalties, non-Euclidean multi-horizon surrogates (e.g., on graph domains), and further integration with uncertainty-calibrated policies for high-stakes decision making. Extensions to variable-density, multiphase, or hybrid physical/digital twin systems remain open areas for demonstration of multi-horizon surrogate generality.

Source: https://www.emergentmind.com/topics/multi-horizon-surrogates