Nash-Sutcliffe Linear Regression
- Nash-Sutcliffe linear regression is a multi-dimensional linear model that defines its estimator by minimizing the Nash-Sutcliffe loss, a weighted loss emphasizing lower internal variance.
- It incorporates data-dependent weights and a decision-theoretic foundation to ensure strict consistency and unbiased estimation of the Nash-Sutcliffe functional.
- Empirical tests using simulations and hydrometeorological data demonstrate that this approach yields significant forecasting improvements compared to traditional OLS methods.
Nash-Sutcliffe linear regression is a multi-dimensional linear model estimated by minimizing the average Nash-Sutcliffe loss, the negatively oriented counterpart of the Nash-Sutcliffe efficiency. In this formulation, the usual forecast-evaluation score is given a decision-theoretic interpretation: the loss is strictly consistent for an elicitable and identifiable multi-dimensional functional, called the Nash-Sutcliffe functional, and the resulting estimator reduces to a data-weighted least squares problem (Tyralis et al., 1 Mar 2026). The framework is designed for forecasting and evaluating multiple time series, and it is explicitly motivated by the observation that maximizing the average NSE across multiple series is the sample analog of minimizing the expected Nash-Sutcliffe loss under a specific stochastic interpretation (Tyralis et al., 1 Mar 2026).
1. Definition through NSE and Nash-Sutcliffe loss
For a one-dimensional time series of length , Nash-Sutcliffe efficiency is defined as
where is the -vector of ones and is the sample mean of the observation vector (Tyralis et al., 1 Mar 2026). The loss-oriented version is
with
and
In the expanded form used throughout the theory,
For multiple time series arranged as a matrix 0, the realized Nash-Sutcliffe loss is
1
and satisfies the identity
2
The paper also introduces the extended loss
3
with corresponding weight
4
to address the case in which the denominator can vanish or become numerically unstable (Tyralis et al., 1 Mar 2026).
2. Decision-theoretic foundation and the Nash-Sutcliffe functional
The central theoretical result is that 5 is strictly consistent for a multi-dimensional functional derived from a weighted-loss construction (Tyralis et al., 1 Mar 2026). The base loss 6 is strictly consistent for the component-wise mean
7
Applying the weighted-loss theorem yields the elicited functional
8
This functional is named the Nash-Sutcliffe functional (Tyralis et al., 1 Mar 2026).
The authors characterize it as a data-weighted component-wise mean because each coordinate mean is reweighted by the same scalar function of the whole observation vector,
9
The weights favor realizations with lower within-vector variability. The Nash-Sutcliffe functional coincides with the ordinary component-wise mean if and only if
0
or equivalently if the pairs 1 are uncorrelated (Tyralis et al., 1 Mar 2026).
The same paper also establishes identifiability. A strict identification function is
2
with empirical version
3
The empirical Nash-Sutcliffe climatology is the zero of this empirical identification function: 4 with component form
5
A direct convexity argument is also given. The expected loss is
6
its gradient is
7
and the Hessian is diagonal with positive entries
8
Hence the expected loss is strictly convex and has the unique minimizer
9
3. Formal regression model and estimator
The Nash-Sutcliffe linear model in the 0 setting is
1
Its estimator is defined by minimization of realized Nash-Sutcliffe loss: 2 The closed-form solution is
3
4
where
5
and
6
The corresponding predictive model is
7
The same paper gives a parallel forecasting-oriented formulation in the 8 orientation: 9 with
0
The decisive structural point is that minimizing the average 1 becomes a weighted least squares problem, but the weights are data-dependent and determined by each response vector’s internal variability relative to its own mean (Tyralis et al., 1 Mar 2026).
4. Relation to ordinary least squares and data orientation
Classical one-dimensional regression minimizes MSE and yields
2
while multivariate linear regression minimizes the Euclidean norm loss and produces
3
Nash-Sutcliffe regression replaces this unweighted objective with the weighted objective above (Tyralis et al., 1 Mar 2026).
The interpretation of the weights is explicit: Nash-Sutcliffe regression gives more weight to response vectors with smaller internal variance, because these observations make the denominator of the NSE smaller. Ordinary least squares, by contrast, treats all series equally (Tyralis et al., 1 Mar 2026). This difference is not merely algebraic; it changes the target functional from the ordinary component-wise mean to the Nash-Sutcliffe functional.
The paper also emphasizes that the matrix orientation encodes the stochastic interpretation. In the 4 setting, each column 5 is treated as one realization of the same 6-dimensional random vector 7. Under that interpretation, the average NSE across the 8 series corresponds to evaluating predictions of a single underlying random vector. The authors state that this assumption is violated if one mixes time series with different stochastic structures, such as daily and monthly records, streamflow and temperature, or series from different regimes or processes (Tyralis et al., 1 Mar 2026).
In Section 5, the matrix is transposed and each row is interpreted as a realization of a 9-dimensional vector across variables or series at a fixed time. The corresponding realized loss is
0
The paper describes this as the more natural setup for forecasting. It also notes a specific result in Section 5.6: if one estimates the forecasting model and then evaluates with the realized NSE in the 1 setup, the optimization decomposes in a way that recovers the same coefficient estimates as standard multivariate least squares, because the weights are constant with respect to each row-specific parameter block (Tyralis et al., 1 Mar 2026).
5. Empirical behavior in simulations and hydrometeorological forecasting
The empirical section contains three simulation studies and two real-data applications (Tyralis et al., 1 Mar 2026). In Simulation 1, 2 series of length 3 are generated under five scenarios, including Gaussian, log-normal, independent, dependent, and heterogeneous mean structures. For IID Gaussian data, the component-wise mean climatology and the Nash-Sutcliffe climatology are nearly identical; for log-normal and dependent non-Gaussian settings, they diverge substantially; and in the most non-Gaussian case, the Nash-Sutcliffe climatology has much better realized NSE than the component-wise mean climatology.
In Simulation 2, the authors simulate a linear model with 4, 5, 6, and correlated log-normal errors in the 7 setting. OLS performs better under Euclidean norm loss, while Nash-Sutcliffe regression performs dramatically better under realized Nash-Sutcliffe loss. The reported test-set values are Euclidean norm loss 8 for OLS and 9 for Nash-Sutcliffe regression, and Nash-Sutcliffe loss 0 for OLS and 1 for Nash-Sutcliffe regression (Tyralis et al., 1 Mar 2026).
In Simulation 3, the same data-generating process is transposed to the 2 forecasting setting. The pattern remains the same: OLS is best for Euclidean norm loss, while Nash-Sutcliffe regression is best for realized Nash-Sutcliffe loss. The reported test-set values are Euclidean norm loss 3 for OLS and 4 for Nash-Sutcliffe regression, and Nash-Sutcliffe loss 5 for OLS and 6 for Nash-Sutcliffe regression (Tyralis et al., 1 Mar 2026).
The real-data applications use streamflow and temperature data from ten French river basins. The following test-set results are reported.
| Setting | Euclidean norm loss | Nash-Sutcliffe loss |
|---|---|---|
| Streamflow, one-dimensional regression | 2.6781 | 0.3791 |
| Streamflow, multivariate regression | 2.5359 | 0.2244 |
| Streamflow, Nash-Sutcliffe regression | 2.6214 | 0.1222 |
| Temperature, one-dimensional regression | 43.8880 | 3.5512 |
| Temperature, multivariate regression | 32.6507 | 2.4006 |
| Temperature, Nash-Sutcliffe regression | 34.7666 | 2.2500 |
For streamflow, Nash-Sutcliffe regression gives the best NSE-based performance, with about a 68% reduction relative to one-dimensional regression and 46% relative to multivariate regression. For temperature, the gains are smaller, which the authors attribute to temperature being closer to Gaussian, so the Nash-Sutcliffe functional is nearer to the component-wise mean (Tyralis et al., 1 Mar 2026).
6. Scope, limitations, and terminological boundaries
The framework is explicit about several limitations. First, because
7
the weight is infinite when a series has zero internal variance, so the strict-consistency results require the support of the distribution to exclude that set, or else one uses the extended loss 8 with 9 (Tyralis et al., 1 Mar 2026). Second, averaging NSE across series is only coherent under the assumption that the series can be viewed as realizations of the same stochastic process. Third, training must match evaluation: if model evaluation uses NSE, the model should be trained using the Nash-Sutcliffe loss, not MSE, unless special conditions hold. Fourth, the 0 and 1 formulations are not merely notationally different; they encode different assumptions about whether the stochastic process is replicated across series or across time. The same paper adds that, for forecasting multiple time series, global models are often more appropriate than fitting separate local models (Tyralis et al., 1 Mar 2026).
The term also benefits from terminological separation from other uses of “Nash” in the regression literature. In "Estimating strength of DDoS attack using various regression models" (Gupta et al., 2012), the Nash-Sutcliffe efficiency index 2 appears as one of several performance measures—alongside 3, CC, SSE, MSE, RMSE, NMSE, and MAE—for comparing linear, polynomial, logarithmic, power, and exponential regression models. In that setting, polynomial regression performs best overall, and 4 is reported to quantify predictive quality rather than to define the estimator itself (Gupta et al., 2012). This contrasts with Nash-Sutcliffe linear regression, where the loss 5 is the estimation criterion (Tyralis et al., 1 Mar 2026).
A further distinction concerns work on strategic behavior in linear regression. "The Effect of Strategic Noise in Linear Regression" studies 6-norm minimization with convex regularization, proves existence of a pure Nash equilibrium for 7, establishes a unique pure Nash equilibrium outcome, and analyzes equilibrium quality through the pure price of anarchy (Hossain et al., 2020). Here “Nash” refers to Nash equilibrium. By contrast, Nash-Sutcliffe linear regression concerns the Nash-Sutcliffe efficiency and its associated loss. Likewise, "Almost linear Nash groups" studies Nash manifolds, Nash groups, and Nash representations with finite kernel; there “Nash” refers to Nash geometry rather than regression loss or strategic equilibrium (Sun, 2013).