Nash–Sutcliffe Loss: A Decision-Theoretic Approach
- Nash–Sutcliffe loss is a normalized squared-error metric that reformulates NSE, comparing model predictions against a climatological benchmark to assess performance.
- The loss is strictly consistent and elicitable for a weighted component-wise mean, where inverse within-series variability determines the target functional.
- Extensions to multi-series settings and linear regression highlight differences from ordinary least squares, emphasizing decision‐theoretic evaluation in hydrology.
Searching arXiv for the relevant Nash–Sutcliffe loss papers to ground the article in current research. Nash–Sutcliffe loss is the negatively oriented counterpart of Nash–Sutcliffe efficiency (NSE), defined by the identity . For a single time series of length , with observations and predictions , it takes the normalized squared-error form
where is the sample mean of and is the -vector of ones (Tyralis et al., 1 Mar 2026). In hydrology and environmental modeling, NSE is a standard positively oriented skill score with range , whereas Nash–Sutcliffe loss is minimized rather than maximized. Recent work places this loss within a decision-theoretic framework, showing that its expected value is minimized not by the ordinary mean in general, but by a distinct weighted target functional across multiple time series (Tyralis et al., 1 Mar 2026). Closely related research on normalized error objectives further clarifies how Nash–Sutcliffe loss differs from other bounded agreement-style criteria such as the negatively oriented index of agreement loss 0 and its variant 1 (Tyralis et al., 16 Oct 2025).
1. Definition and basic relation to Nash–Sutcliffe efficiency
For predictions 2 and observations 3, NSE is defined as
4
Expanding the denominator yields
5
The corresponding negatively oriented loss is
6
and equivalently
7
This representation makes explicit that Nash–Sutcliffe loss is a normalized Euclidean squared-error criterion. The numerator is the ordinary Euclidean norm loss
8
while the denominator is the squared error of the climatological benchmark that always predicts the sample mean of the same realization (Tyralis et al., 1 Mar 2026).
For a single time series, this normalization does not alter model rankings relative to MSE, because the denominator is fixed with respect to model parameters on the training sample. A related linear-regression analysis states that, for evaluations involving a single time series, “9 and 0 lead to identical parameter estimates when fitting a semiparametric regression model to a single time series” (Tyralis et al., 8 Jun 2026). This establishes the single-series equivalence between maximizing NSE and minimizing normalized MSE.
2. Multi-series formulation and the meaning of averaging
A central setting for Nash–Sutcliffe loss is evaluation across multiple time series. With 1 time series of common length 2, observations are arranged in a matrix 3, where each column 4 is one series, and predictions are collected in 5. The common empirical practice is to compute NSE for each series and average: 6 Equivalently,
7
The realized average Nash–Sutcliffe loss is therefore
8
The decision-theoretic interpretation of this average is precise. The average loss is the sample analog of minimizing
9
under a joint distribution 0 for the random vector 1 (Tyralis et al., 1 Mar 2026). The cited paper interprets this to mean that averaging NSE across columns is statistically coherent only if the columns 2 are realizations of the same 3-dimensional random vector. The explicit formulation is that common averaging of NSE across series “implicitly assumes that all series originate from a single non-stationary stochastic process” (Tyralis et al., 1 Mar 2026).
This point is consequential for empirical design. In the 4 arrangement, each column is a full series, each row is a time index, and the interpretation treats the series as replicates of one 5-dimensional stochastic object. A plausible implication is that indiscriminate aggregation across heterogeneous series types lacks the population-level justification provided for the common-distribution setting.
3. Decision-theoretic status: strict consistency, elicitability, and identifiability
The main theoretical advance is the proof that Nash–Sutcliffe loss is strictly consistent for a specific elicitable and identifiable functional, termed the Nash–Sutcliffe functional (Tyralis et al., 1 Mar 2026). The functional is defined as
6
with weight
7
Thus the target is a weighted version of the component-wise mean rather than the ordinary component-wise mean (Tyralis et al., 1 Mar 2026).
The strict-consistency result is established on a class 8 for which 9 is integrable and the corresponding weighted distribution remains in 0. Under that assumption,
1
is strictly 2-consistent for 3 (Tyralis et al., 1 Mar 2026).
The same work gives the direct minimization argument. The expected loss is
4
Its derivatives are
5
and the second derivatives satisfy
6
Hence the expected loss is strictly convex, and its unique minimizer is given componentwise by
7
The same paper also proves identifiability. A strict 8-identification function is
9
with empirical version
0
The theoretical significance is that Nash–Sutcliffe loss is not merely a heuristic score written in loss form. It is a proper loss for a well-defined multi-dimensional target functional. This distinguishes the recent decision-theoretic treatment from earlier use of NSE purely as a performance summary.
4. The Nash–Sutcliffe functional and its difference from the mean
The formula
1
implies that minimizing expected Nash–Sutcliffe loss yields a data-weighted component-wise mean (Tyralis et al., 1 Mar 2026). The weight
2
depends on the within-series variability of the realization around its own sample mean. The cited paper explicitly interprets this as assigning larger weight to vectors with small internal variability and smaller weight to vectors with large internal variability (Tyralis et al., 1 Mar 2026).
This target differs from the ordinary component-wise mean
3
which is elicited by the Euclidean norm loss 4 (Tyralis et al., 1 Mar 2026). The two functionals coincide only under
5
equivalently
6
Componentwise, this requires each pair 7 to be uncorrelated. The paper states that this is rare in practice, and identifies one special case where equality holds: 8 with 9 (Tyralis et al., 1 Mar 2026).
Simulation evidence in the same work reflects this distinction. Under IID Gaussian data, the ordinary component-wise mean climatology and the Nash–Sutcliffe climatology are nearly identical, whereas under log-normal, dependent, or heterogeneous settings they differ substantially (Tyralis et al., 1 Mar 2026). This supports the theoretical claim that NSE-based estimation generally targets something different from least squares.
A closely related comparison appears in research on normalized agreement losses. There, the quantity 0 is described as a normalized squared-error loss with denominator 1, contrasting with the index-of-agreement-based loss 2, which uses the same squared-error numerator but a different denominator involving both predictions and observations relative to the mean (Tyralis et al., 16 Oct 2025). That comparison further situates Nash–Sutcliffe loss as a variance-relative skill criterion rather than an agreement-style bounded ratio.
5. Estimation, regression, and data orientation
For a constant prediction vector, the sample analog of the Nash–Sutcliffe functional is the weighted climatology
3
and the associated 4-estimator is
5
whose solution is the same weighted average (Tyralis et al., 1 Mar 2026).
The linear-model extension is Nash–Sutcliffe linear regression. For response 6 and predictor vector 7, the model is
8
Given response matrix 9 and predictor matrix 0, the estimator is
1
Its closed-form solution is
2
with
3
and
4
This is explicitly a data-weighted least squares estimator. Unlike multivariate OLS,
5
Nash–Sutcliffe regression weights series by inverse within-series variability (Tyralis et al., 1 Mar 2026).
The same paper also proposes a different orientation for forecasting multiple stationary dependent time series. Instead of 6, it uses 7, where each row 8 is one realization of a 9-dimensional random vector and 0 is the number of time points. The row-wise realized loss is
1
The authors describe this reorientation as “a more natural empirical implementation” for forecasting multiple stationary dependent time series with differing stochastic properties (Tyralis et al., 1 Mar 2026).
An important subtlety is that when the row-wise forecasting linear model is evaluated using the original column-wise realized loss, the minimization separates into 2 independent weighted least squares problems with the same minimizers as unweighted least squares, reducing the estimator to standard multivariate OLS (Tyralis et al., 1 Mar 2026). This suggests that the empirical meaning of Nash–Sutcliffe optimization depends materially on how series and time are arranged.
6. Extensions, related losses, and contrasts with alternative objectives
To handle zero denominators, the extended loss
3
is introduced (Tyralis et al., 1 Mar 2026). It is strictly consistent for the extended Nash–Sutcliffe functional
4
The explicit statement is that adding 5 changes the target functional (Tyralis et al., 1 Mar 2026).
The denominator pathology motivating this extension is that
6
diverges when all components of 7 are equal. The theory is therefore developed on distribution classes for which 8 is integrable, the corresponding weighted distribution remains in 9, and 0 (Tyralis et al., 1 Mar 2026).
Nash–Sutcliffe loss also sits within a broader class of normalized error objectives. The study of the index of agreement loss 1 makes the comparison explicit by writing
2
and emphasizing that NSE and 3 share the same squared-error numerator but differ in denominator geometry (Tyralis et al., 16 Oct 2025). In that framework, 4 and 5 are always bounded in 6, while 7 and exceeds 8 when the model is worse than mean climatology (Tyralis et al., 16 Oct 2025). The comparison is important because it shows that not all normalized squared-error ratios are inferentially equivalent.
A further contrast arises with Kling–Gupta efficiency (KGE). In single-series regression, one analysis states that maximizing NSE is equivalent to minimizing MSE and hence to ordinary least squares, while maximizing KGE is equivalent to minimizing the negatively oriented Kling–Gupta loss
9
(Tyralis et al., 8 Jun 2026). The paper proves that no single estimator can generally maximize both NSE and KGE simultaneously. In its linear-regression setting, OLS maximizes NSE, whereas Kling–Gupta regression minimizes 00 by scaling the OLS coefficient vector with a variance-inflation factor so that fitted values match the sample variance of the response on the training set (Tyralis et al., 8 Jun 2026). This identifies a structural difference between NSE-oriented estimation, which favors variance reduction, and KGE-oriented estimation, which favors variance replication.
7. Empirical behavior and practical interpretation
Simulation and application results indicate that the distinction between MSE and Nash–Sutcliffe loss is empirically consequential when the target functional differs materially from the mean. Under non-Gaussian correlated log-normal errors, Nash–Sutcliffe linear regression “slightly worsens Euclidean norm loss but massively improves realized Nash-Sutcliffe loss compared with ordinary multivariate regression” (Tyralis et al., 1 Mar 2026). In the forecasting-oriented 01 setup, the same qualitative pattern persists: Nash–Sutcliffe regression dominates on realized 02, whereas OLS dominates on Euclidean norm loss (Tyralis et al., 1 Mar 2026).
Real-data hydrology applications using daily streamflow and temperature from 10 French river basins reinforce this alignment principle. For streamflow, multivariate OLS has the best Euclidean norm loss, whereas Nash–Sutcliffe regression has the best Nash–Sutcliffe loss by a large margin. For temperature, Nash–Sutcliffe regression again performs best under Nash–Sutcliffe loss, but the gains are smaller; the paper suggests that temperature may be closer to Gaussian, bringing the Nash–Sutcliffe functional closer to the ordinary mean (Tyralis et al., 1 Mar 2026). This suggests that the practical gap between mean-oriented and Nash–Sutcliffe-oriented training depends on the data-generating structure.
Related evidence from agreement-ratio losses illustrates the boundary of this conclusion. In hydrologic calibration of the GR4J daily lumped hydrologic model on 10 French catchments, squared error, 03, and 04 produced highly similar results, especially when predictor–response dependence was strong (Tyralis et al., 16 Oct 2025). That study interprets the similarity as analogous to high-correlation linear regression, where squared-error fitting and some normalized agreement losses converge in behavior. However, it also shows theoretically that 05 and 06 do not uniquely elicit the mean in scalar fitting, whereas squared error and Nash–Sutcliffe-type criteria remain tied to mean-oriented targets in the single-series setting (Tyralis et al., 16 Oct 2025).
The resulting interpretation of Nash–Sutcliffe loss is therefore specific rather than generic. It is not simply “bounded NSE,” nor is it interchangeable with index-of-agreement losses or KGE-based objectives. Its defining property is that it rewrites NSE as a loss suitable for estimation and evaluation, and recent theory shows that, in the multi-series formulation, it is strictly consistent for a weighted component-wise mean determined by inverse within-series variability (Tyralis et al., 1 Mar 2026). This decision-theoretic grounding clarifies both the appeal of NSE-based learning and the limits of treating NSE as merely a normalized form of MSE.