Papers
Topics
Authors
Recent
Search
2000 character limit reached

Nash-Sutcliffe Linear Regression

Updated 5 July 2026
  • Nash-Sutcliffe linear regression is a multi-dimensional linear model that defines its estimator by minimizing the Nash-Sutcliffe loss, a weighted loss emphasizing lower internal variance.
  • It incorporates data-dependent weights and a decision-theoretic foundation to ensure strict consistency and unbiased estimation of the Nash-Sutcliffe functional.
  • Empirical tests using simulations and hydrometeorological data demonstrate that this approach yields significant forecasting improvements compared to traditional OLS methods.

Nash-Sutcliffe linear regression is a multi-dimensional linear model estimated by minimizing the average Nash-Sutcliffe loss, the negatively oriented counterpart of the Nash-Sutcliffe efficiency. In this formulation, the usual forecast-evaluation score is given a decision-theoretic interpretation: the loss is strictly consistent for an elicitable and identifiable multi-dimensional functional, called the Nash-Sutcliffe functional, and the resulting estimator reduces to a data-weighted least squares problem (Tyralis et al., 1 Mar 2026). The framework is designed for forecasting and evaluating multiple time series, and it is explicitly motivated by the observation that maximizing the average NSE across multiple series is the sample analog of minimizing the expected Nash-Sutcliffe loss under a specific stochastic interpretation (Tyralis et al., 1 Mar 2026).

1. Definition through NSE and Nash-Sutcliffe loss

For a one-dimensional time series of length d2d \ge 2, Nash-Sutcliffe efficiency is defined as

NSE(Zd,Yd):=1MSE(Zd,yd)/MSE(1du(yd),yd),d2,\mathrm{NSE}(Z_d,Y_d) := 1-\mathrm{MSE}(Z_d,y_d)/\mathrm{MSE}(1_d u(y_d), y_d), \quad d \ge 2,

where 1d1_d is the dd-vector of ones and u(yd)u(y_d) is the sample mean of the observation vector (Tyralis et al., 1 Mar 2026). The loss-oriented version is

LNS(Zd,yd)=1NSE(Zd,yd)=w(yd)LEN(Zd,yd),d2,L_{\mathrm{NS}}(Z_d,y_d)=1-\mathrm{NSE}(Z_d,y_d)=w(y_d)L_{\mathrm{EN}}(Z_d,y_d), \quad d \ge 2,

with

LEN(Zd,yd)=zdyd22L_{\mathrm{EN}}(Z_d,y_d)=\|z_d-y_d\|_2^2

and

w(yd)=1/u(yd)1dyd22.w(y_d)=1/\|u(y_d)1_d-y_d\|_2^2.

In the expanded form used throughout the theory,

LNS(zd,yd)=zdyd22u(yd)1dyd22,d2.L_{\mathrm{NS}}(z_d,y_d)=\frac{\|z_d-y_d\|_2^2}{\|u(y_d)1_d-y_d\|_2^2}, \quad d\ge 2.

For multiple time series arranged as a d×nd\times n matrix NSE(Zd,Yd):=1MSE(Zd,yd)/MSE(1du(yd),yd),d2,\mathrm{NSE}(Z_d,Y_d) := 1-\mathrm{MSE}(Z_d,y_d)/\mathrm{MSE}(1_d u(y_d), y_d), \quad d \ge 2,0, the realized Nash-Sutcliffe loss is

NSE(Zd,Yd):=1MSE(Zd,yd)/MSE(1du(yd),yd),d2,\mathrm{NSE}(Z_d,Y_d) := 1-\mathrm{MSE}(Z_d,y_d)/\mathrm{MSE}(1_d u(y_d), y_d), \quad d \ge 2,1

and satisfies the identity

NSE(Zd,Yd):=1MSE(Zd,yd)/MSE(1du(yd),yd),d2,\mathrm{NSE}(Z_d,Y_d) := 1-\mathrm{MSE}(Z_d,y_d)/\mathrm{MSE}(1_d u(y_d), y_d), \quad d \ge 2,2

The paper also introduces the extended loss

NSE(Zd,Yd):=1MSE(Zd,yd)/MSE(1du(yd),yd),d2,\mathrm{NSE}(Z_d,Y_d) := 1-\mathrm{MSE}(Z_d,y_d)/\mathrm{MSE}(1_d u(y_d), y_d), \quad d \ge 2,3

with corresponding weight

NSE(Zd,Yd):=1MSE(Zd,yd)/MSE(1du(yd),yd),d2,\mathrm{NSE}(Z_d,Y_d) := 1-\mathrm{MSE}(Z_d,y_d)/\mathrm{MSE}(1_d u(y_d), y_d), \quad d \ge 2,4

to address the case in which the denominator can vanish or become numerically unstable (Tyralis et al., 1 Mar 2026).

2. Decision-theoretic foundation and the Nash-Sutcliffe functional

The central theoretical result is that NSE(Zd,Yd):=1MSE(Zd,yd)/MSE(1du(yd),yd),d2,\mathrm{NSE}(Z_d,Y_d) := 1-\mathrm{MSE}(Z_d,y_d)/\mathrm{MSE}(1_d u(y_d), y_d), \quad d \ge 2,5 is strictly consistent for a multi-dimensional functional derived from a weighted-loss construction (Tyralis et al., 1 Mar 2026). The base loss NSE(Zd,Yd):=1MSE(Zd,yd)/MSE(1du(yd),yd),d2,\mathrm{NSE}(Z_d,Y_d) := 1-\mathrm{MSE}(Z_d,y_d)/\mathrm{MSE}(1_d u(y_d), y_d), \quad d \ge 2,6 is strictly consistent for the component-wise mean

NSE(Zd,Yd):=1MSE(Zd,yd)/MSE(1du(yd),yd),d2,\mathrm{NSE}(Z_d,Y_d) := 1-\mathrm{MSE}(Z_d,y_d)/\mathrm{MSE}(1_d u(y_d), y_d), \quad d \ge 2,7

Applying the weighted-loss theorem yields the elicited functional

NSE(Zd,Yd):=1MSE(Zd,yd)/MSE(1du(yd),yd),d2,\mathrm{NSE}(Z_d,Y_d) := 1-\mathrm{MSE}(Z_d,y_d)/\mathrm{MSE}(1_d u(y_d), y_d), \quad d \ge 2,8

This functional is named the Nash-Sutcliffe functional (Tyralis et al., 1 Mar 2026).

The authors characterize it as a data-weighted component-wise mean because each coordinate mean is reweighted by the same scalar function of the whole observation vector,

NSE(Zd,Yd):=1MSE(Zd,yd)/MSE(1du(yd),yd),d2,\mathrm{NSE}(Z_d,Y_d) := 1-\mathrm{MSE}(Z_d,y_d)/\mathrm{MSE}(1_d u(y_d), y_d), \quad d \ge 2,9

The weights favor realizations with lower within-vector variability. The Nash-Sutcliffe functional coincides with the ordinary component-wise mean if and only if

1d1_d0

or equivalently if the pairs 1d1_d1 are uncorrelated (Tyralis et al., 1 Mar 2026).

The same paper also establishes identifiability. A strict identification function is

1d1_d2

with empirical version

1d1_d3

The empirical Nash-Sutcliffe climatology is the zero of this empirical identification function: 1d1_d4 with component form

1d1_d5

A direct convexity argument is also given. The expected loss is

1d1_d6

its gradient is

1d1_d7

and the Hessian is diagonal with positive entries

1d1_d8

Hence the expected loss is strictly convex and has the unique minimizer

1d1_d9

(Tyralis et al., 1 Mar 2026).

3. Formal regression model and estimator

The Nash-Sutcliffe linear model in the dd0 setting is

dd1

Its estimator is defined by minimization of realized Nash-Sutcliffe loss: dd2 The closed-form solution is

dd3

dd4

where

dd5

and

dd6

The corresponding predictive model is

dd7

The same paper gives a parallel forecasting-oriented formulation in the dd8 orientation: dd9 with

u(yd)u(y_d)0

The decisive structural point is that minimizing the average u(yd)u(y_d)1 becomes a weighted least squares problem, but the weights are data-dependent and determined by each response vector’s internal variability relative to its own mean (Tyralis et al., 1 Mar 2026).

4. Relation to ordinary least squares and data orientation

Classical one-dimensional regression minimizes MSE and yields

u(yd)u(y_d)2

while multivariate linear regression minimizes the Euclidean norm loss and produces

u(yd)u(y_d)3

Nash-Sutcliffe regression replaces this unweighted objective with the weighted objective above (Tyralis et al., 1 Mar 2026).

The interpretation of the weights is explicit: Nash-Sutcliffe regression gives more weight to response vectors with smaller internal variance, because these observations make the denominator of the NSE smaller. Ordinary least squares, by contrast, treats all series equally (Tyralis et al., 1 Mar 2026). This difference is not merely algebraic; it changes the target functional from the ordinary component-wise mean to the Nash-Sutcliffe functional.

The paper also emphasizes that the matrix orientation encodes the stochastic interpretation. In the u(yd)u(y_d)4 setting, each column u(yd)u(y_d)5 is treated as one realization of the same u(yd)u(y_d)6-dimensional random vector u(yd)u(y_d)7. Under that interpretation, the average NSE across the u(yd)u(y_d)8 series corresponds to evaluating predictions of a single underlying random vector. The authors state that this assumption is violated if one mixes time series with different stochastic structures, such as daily and monthly records, streamflow and temperature, or series from different regimes or processes (Tyralis et al., 1 Mar 2026).

In Section 5, the matrix is transposed and each row is interpreted as a realization of a u(yd)u(y_d)9-dimensional vector across variables or series at a fixed time. The corresponding realized loss is

LNS(Zd,yd)=1NSE(Zd,yd)=w(yd)LEN(Zd,yd),d2,L_{\mathrm{NS}}(Z_d,y_d)=1-\mathrm{NSE}(Z_d,y_d)=w(y_d)L_{\mathrm{EN}}(Z_d,y_d), \quad d \ge 2,0

The paper describes this as the more natural setup for forecasting. It also notes a specific result in Section 5.6: if one estimates the forecasting model and then evaluates with the realized NSE in the LNS(Zd,yd)=1NSE(Zd,yd)=w(yd)LEN(Zd,yd),d2,L_{\mathrm{NS}}(Z_d,y_d)=1-\mathrm{NSE}(Z_d,y_d)=w(y_d)L_{\mathrm{EN}}(Z_d,y_d), \quad d \ge 2,1 setup, the optimization decomposes in a way that recovers the same coefficient estimates as standard multivariate least squares, because the weights are constant with respect to each row-specific parameter block (Tyralis et al., 1 Mar 2026).

5. Empirical behavior in simulations and hydrometeorological forecasting

The empirical section contains three simulation studies and two real-data applications (Tyralis et al., 1 Mar 2026). In Simulation 1, LNS(Zd,yd)=1NSE(Zd,yd)=w(yd)LEN(Zd,yd),d2,L_{\mathrm{NS}}(Z_d,y_d)=1-\mathrm{NSE}(Z_d,y_d)=w(y_d)L_{\mathrm{EN}}(Z_d,y_d), \quad d \ge 2,2 series of length LNS(Zd,yd)=1NSE(Zd,yd)=w(yd)LEN(Zd,yd),d2,L_{\mathrm{NS}}(Z_d,y_d)=1-\mathrm{NSE}(Z_d,y_d)=w(y_d)L_{\mathrm{EN}}(Z_d,y_d), \quad d \ge 2,3 are generated under five scenarios, including Gaussian, log-normal, independent, dependent, and heterogeneous mean structures. For IID Gaussian data, the component-wise mean climatology and the Nash-Sutcliffe climatology are nearly identical; for log-normal and dependent non-Gaussian settings, they diverge substantially; and in the most non-Gaussian case, the Nash-Sutcliffe climatology has much better realized NSE than the component-wise mean climatology.

In Simulation 2, the authors simulate a linear model with LNS(Zd,yd)=1NSE(Zd,yd)=w(yd)LEN(Zd,yd),d2,L_{\mathrm{NS}}(Z_d,y_d)=1-\mathrm{NSE}(Z_d,y_d)=w(y_d)L_{\mathrm{EN}}(Z_d,y_d), \quad d \ge 2,4, LNS(Zd,yd)=1NSE(Zd,yd)=w(yd)LEN(Zd,yd),d2,L_{\mathrm{NS}}(Z_d,y_d)=1-\mathrm{NSE}(Z_d,y_d)=w(y_d)L_{\mathrm{EN}}(Z_d,y_d), \quad d \ge 2,5, LNS(Zd,yd)=1NSE(Zd,yd)=w(yd)LEN(Zd,yd),d2,L_{\mathrm{NS}}(Z_d,y_d)=1-\mathrm{NSE}(Z_d,y_d)=w(y_d)L_{\mathrm{EN}}(Z_d,y_d), \quad d \ge 2,6, and correlated log-normal errors in the LNS(Zd,yd)=1NSE(Zd,yd)=w(yd)LEN(Zd,yd),d2,L_{\mathrm{NS}}(Z_d,y_d)=1-\mathrm{NSE}(Z_d,y_d)=w(y_d)L_{\mathrm{EN}}(Z_d,y_d), \quad d \ge 2,7 setting. OLS performs better under Euclidean norm loss, while Nash-Sutcliffe regression performs dramatically better under realized Nash-Sutcliffe loss. The reported test-set values are Euclidean norm loss LNS(Zd,yd)=1NSE(Zd,yd)=w(yd)LEN(Zd,yd),d2,L_{\mathrm{NS}}(Z_d,y_d)=1-\mathrm{NSE}(Z_d,y_d)=w(y_d)L_{\mathrm{EN}}(Z_d,y_d), \quad d \ge 2,8 for OLS and LNS(Zd,yd)=1NSE(Zd,yd)=w(yd)LEN(Zd,yd),d2,L_{\mathrm{NS}}(Z_d,y_d)=1-\mathrm{NSE}(Z_d,y_d)=w(y_d)L_{\mathrm{EN}}(Z_d,y_d), \quad d \ge 2,9 for Nash-Sutcliffe regression, and Nash-Sutcliffe loss LEN(Zd,yd)=zdyd22L_{\mathrm{EN}}(Z_d,y_d)=\|z_d-y_d\|_2^20 for OLS and LEN(Zd,yd)=zdyd22L_{\mathrm{EN}}(Z_d,y_d)=\|z_d-y_d\|_2^21 for Nash-Sutcliffe regression (Tyralis et al., 1 Mar 2026).

In Simulation 3, the same data-generating process is transposed to the LEN(Zd,yd)=zdyd22L_{\mathrm{EN}}(Z_d,y_d)=\|z_d-y_d\|_2^22 forecasting setting. The pattern remains the same: OLS is best for Euclidean norm loss, while Nash-Sutcliffe regression is best for realized Nash-Sutcliffe loss. The reported test-set values are Euclidean norm loss LEN(Zd,yd)=zdyd22L_{\mathrm{EN}}(Z_d,y_d)=\|z_d-y_d\|_2^23 for OLS and LEN(Zd,yd)=zdyd22L_{\mathrm{EN}}(Z_d,y_d)=\|z_d-y_d\|_2^24 for Nash-Sutcliffe regression, and Nash-Sutcliffe loss LEN(Zd,yd)=zdyd22L_{\mathrm{EN}}(Z_d,y_d)=\|z_d-y_d\|_2^25 for OLS and LEN(Zd,yd)=zdyd22L_{\mathrm{EN}}(Z_d,y_d)=\|z_d-y_d\|_2^26 for Nash-Sutcliffe regression (Tyralis et al., 1 Mar 2026).

The real-data applications use streamflow and temperature data from ten French river basins. The following test-set results are reported.

Setting Euclidean norm loss Nash-Sutcliffe loss
Streamflow, one-dimensional regression 2.6781 0.3791
Streamflow, multivariate regression 2.5359 0.2244
Streamflow, Nash-Sutcliffe regression 2.6214 0.1222
Temperature, one-dimensional regression 43.8880 3.5512
Temperature, multivariate regression 32.6507 2.4006
Temperature, Nash-Sutcliffe regression 34.7666 2.2500

For streamflow, Nash-Sutcliffe regression gives the best NSE-based performance, with about a 68% reduction relative to one-dimensional regression and 46% relative to multivariate regression. For temperature, the gains are smaller, which the authors attribute to temperature being closer to Gaussian, so the Nash-Sutcliffe functional is nearer to the component-wise mean (Tyralis et al., 1 Mar 2026).

6. Scope, limitations, and terminological boundaries

The framework is explicit about several limitations. First, because

LEN(Zd,yd)=zdyd22L_{\mathrm{EN}}(Z_d,y_d)=\|z_d-y_d\|_2^27

the weight is infinite when a series has zero internal variance, so the strict-consistency results require the support of the distribution to exclude that set, or else one uses the extended loss LEN(Zd,yd)=zdyd22L_{\mathrm{EN}}(Z_d,y_d)=\|z_d-y_d\|_2^28 with LEN(Zd,yd)=zdyd22L_{\mathrm{EN}}(Z_d,y_d)=\|z_d-y_d\|_2^29 (Tyralis et al., 1 Mar 2026). Second, averaging NSE across series is only coherent under the assumption that the series can be viewed as realizations of the same stochastic process. Third, training must match evaluation: if model evaluation uses NSE, the model should be trained using the Nash-Sutcliffe loss, not MSE, unless special conditions hold. Fourth, the w(yd)=1/u(yd)1dyd22.w(y_d)=1/\|u(y_d)1_d-y_d\|_2^2.0 and w(yd)=1/u(yd)1dyd22.w(y_d)=1/\|u(y_d)1_d-y_d\|_2^2.1 formulations are not merely notationally different; they encode different assumptions about whether the stochastic process is replicated across series or across time. The same paper adds that, for forecasting multiple time series, global models are often more appropriate than fitting separate local models (Tyralis et al., 1 Mar 2026).

The term also benefits from terminological separation from other uses of “Nash” in the regression literature. In "Estimating strength of DDoS attack using various regression models" (Gupta et al., 2012), the Nash-Sutcliffe efficiency index w(yd)=1/u(yd)1dyd22.w(y_d)=1/\|u(y_d)1_d-y_d\|_2^2.2 appears as one of several performance measures—alongside w(yd)=1/u(yd)1dyd22.w(y_d)=1/\|u(y_d)1_d-y_d\|_2^2.3, CC, SSE, MSE, RMSE, NMSE, and MAE—for comparing linear, polynomial, logarithmic, power, and exponential regression models. In that setting, polynomial regression performs best overall, and w(yd)=1/u(yd)1dyd22.w(y_d)=1/\|u(y_d)1_d-y_d\|_2^2.4 is reported to quantify predictive quality rather than to define the estimator itself (Gupta et al., 2012). This contrasts with Nash-Sutcliffe linear regression, where the loss w(yd)=1/u(yd)1dyd22.w(y_d)=1/\|u(y_d)1_d-y_d\|_2^2.5 is the estimation criterion (Tyralis et al., 1 Mar 2026).

A further distinction concerns work on strategic behavior in linear regression. "The Effect of Strategic Noise in Linear Regression" studies w(yd)=1/u(yd)1dyd22.w(y_d)=1/\|u(y_d)1_d-y_d\|_2^2.6-norm minimization with convex regularization, proves existence of a pure Nash equilibrium for w(yd)=1/u(yd)1dyd22.w(y_d)=1/\|u(y_d)1_d-y_d\|_2^2.7, establishes a unique pure Nash equilibrium outcome, and analyzes equilibrium quality through the pure price of anarchy (Hossain et al., 2020). Here “Nash” refers to Nash equilibrium. By contrast, Nash-Sutcliffe linear regression concerns the Nash-Sutcliffe efficiency and its associated loss. Likewise, "Almost linear Nash groups" studies Nash manifolds, Nash groups, and Nash representations with finite kernel; there “Nash” refers to Nash geometry rather than regression loss or strategic equilibrium (Sun, 2013).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (4)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Nash-Sutcliffe Linear Regression.