Papers
Topics
Authors
Recent
Search
2000 character limit reached

Nash–Sutcliffe Functional and Loss Framework

Updated 5 July 2026
  • The Nash–Sutcliffe functional is a data-weighted component-wise mean serving as an elicitable target, linking hydrological scoring with a decision-theoretic foundation.
  • It reweights observations via the inverse centered sum-of-squares, ensuring strict consistency and identifiability when optimizing predictive performance.
  • Its framework underpins methods like Nash–Sutcliffe regression, aligning average NSE evaluation with expected loss minimization across multiple series.

Searching arXiv for the cited papers and closely related NSE literature. The Nash–Sutcliffe functional is the elicitable and identifiable multi-dimensional target that is strictly elicited by the Nash–Sutcliffe loss, LNS=1NSEL_{\text{NS}} = 1 - \text{NSE}. In the formulation developed in "Learning with the Nash-Sutcliffe loss," it is a data-weighted component-wise mean, obtained by reweighting the underlying distribution by the inverse centered within-vector sum of squares. This construction gives a decision-theoretic foundation to NSE-based forecast evaluation and model estimation across multiple series, and clarifies the conditions under which maximizing average NSE is the sample analog of minimizing an expected loss (Tyralis et al., 1 Mar 2026).

1. From Nash–Sutcliffe efficiency to Nash–Sutcliffe loss

For a single series ii with observations {yi,t}t=1Ti\{y_{i,t}\}_{t=1}^{T_i}, forecasts {y^i,t}t=1Ti\{\hat y_{i,t}\}_{t=1}^{T_i}, and sample mean

yˉi:=Ti1t=1Tiyi,t,\bar y_i := T_i^{-1} \sum_{t=1}^{T_i} y_{i,t},

the standard sums of squares are

SSEi=t=1Ti(yi,ty^i,t)2,TSSi=t=1Ti(yi,tyˉi)2.\text{SSE}_i = \sum_{t=1}^{T_i} (y_{i,t} - \hat y_{i,t})^2,\qquad \text{TSS}_i = \sum_{t=1}^{T_i} (y_{i,t} - \bar y_i)^2.

The Nash–Sutcliffe efficiency and its negatively oriented counterpart are

NSEi=1SSEiTSSi,LNS,i=1NSEi=SSEiTSSi.\text{NSE}_i = 1 - \frac{\text{SSE}_i}{\text{TSS}_i},\qquad L_{\text{NS},i} = 1 - \text{NSE}_i = \frac{\text{SSE}_i}{\text{TSS}_i}.

In this form, NSEi\text{NSE}_i compares predictive skill to the naïve mean baseline that always predicts yˉi\bar y_i, while LNS,iL_{\text{NS},i} is the corresponding relative squared-error loss. The condition ii0 ensures ii1 (Tyralis et al., 1 Mar 2026).

Across ii2 series, the average Nash–Sutcliffe loss and average NSE are

ii3

In vector notation for a ii4-dimensional realization ii5 and prediction ii6,

ii7

This representation is the bridge from the familiar hydrological score to a population-level scoring rule on ii8 (Tyralis et al., 1 Mar 2026).

Within the hydrological literature, NSE is also characterized as

ii9

Its upper bound is {yi,t}t=1Ti\{y_{i,t}\}_{t=1}^{T_i}0, achieved only for a perfect match, while the lower bound is unbounded. NSE {yi,t}t=1Ti\{y_{i,t}\}_{t=1}^{T_i}1 corresponds to the naïve reference that predicts the mean {yi,t}t=1Ti\{y_{i,t}\}_{t=1}^{T_i}2 at all times; values {yi,t}t=1Ti\{y_{i,t}\}_{t=1}^{T_i}3 indicate performance worse than this baseline, and values {yi,t}t=1Ti\{y_{i,t}\}_{t=1}^{T_i}4 indicate improvement over the baseline (Khatami et al., 2020).

2. Decision-theoretic foundation

The decisive step is to treat the Nash–Sutcliffe loss as a scoring function in a distributional setting. Let {yi,t}t=1Ti\{y_{i,t}\}_{t=1}^{T_i}5 be a {yi,t}t=1Ti\{y_{i,t}\}_{t=1}^{T_i}6-dimensional random vector with joint distribution {yi,t}t=1Ti\{y_{i,t}\}_{t=1}^{T_i}7 on {yi,t}t=1Ti\{y_{i,t}\}_{t=1}^{T_i}8. The population loss is

{yi,t}t=1Ti\{y_{i,t}\}_{t=1}^{T_i}9

This embeds NSE in the scoring-rules framework through a weighted Euclidean loss (Tyralis et al., 1 Mar 2026).

The underlying formal notions are those of elicitability, strict consistency, and identifiability. A functional {y^i,t}t=1Ti\{\hat y_{i,t}\}_{t=1}^{T_i}0 is elicitable if there exists a loss {y^i,t}t=1Ti\{\hat y_{i,t}\}_{t=1}^{T_i}1 such that {y^i,t}t=1Ti\{\hat y_{i,t}\}_{t=1}^{T_i}2 is uniquely minimized at {y^i,t}t=1Ti\{\hat y_{i,t}\}_{t=1}^{T_i}3. A loss is strictly {y^i,t}t=1Ti\{\hat y_{i,t}\}_{t=1}^{T_i}4-consistent for {y^i,t}t=1Ti\{\hat y_{i,t}\}_{t=1}^{T_i}5 if the expected loss is uniquely minimized at the target, and a functional is identifiable if there exists {y^i,t}t=1Ti\{\hat y_{i,t}\}_{t=1}^{T_i}6 such that {y^i,t}t=1Ti\{\hat y_{i,t}\}_{t=1}^{T_i}7 if and only if {y^i,t}t=1Ti\{\hat y_{i,t}\}_{t=1}^{T_i}8 (Tyralis et al., 1 Mar 2026).

The paper invokes Theorem 1 of Gneiting (2011): if {y^i,t}t=1Ti\{\hat y_{i,t}\}_{t=1}^{T_i}9 is a (strictly) consistent loss for an elicitable functional yˉi:=Ti1t=1Tiyi,t,\bar y_i := T_i^{-1} \sum_{t=1}^{T_i} y_{i,t},0, then the weighted loss

yˉi:=Ti1t=1Tiyi,t,\bar y_i := T_i^{-1} \sum_{t=1}^{T_i} y_{i,t},1

is (strictly) consistent for the weighted functional

yˉi:=Ti1t=1Tiyi,t,\bar y_i := T_i^{-1} \sum_{t=1}^{T_i} y_{i,t},2

where yˉi:=Ti1t=1Tiyi,t,\bar y_i := T_i^{-1} \sum_{t=1}^{T_i} y_{i,t},3 has density proportional to yˉi:=Ti1t=1Tiyi,t,\bar y_i := T_i^{-1} \sum_{t=1}^{T_i} y_{i,t},4. Applied with yˉi:=Ti1t=1Tiyi,t,\bar y_i := T_i^{-1} \sum_{t=1}^{T_i} y_{i,t},5 and the Nash–Sutcliffe weight yˉi:=Ti1t=1Tiyi,t,\bar y_i := T_i^{-1} \sum_{t=1}^{T_i} y_{i,t},6, the result yields the Nash–Sutcliffe functional as the target of the Nash–Sutcliffe loss (Tyralis et al., 1 Mar 2026).

3. Definition of the Nash–Sutcliffe functional

The Nash–Sutcliffe functional is the component-wise mean under the reweighted distribution: yˉi:=Ti1t=1Tiyi,t,\bar y_i := T_i^{-1} \sum_{t=1}^{T_i} y_{i,t},7 The paper explicitly describes this as a data-weighted component-wise mean (Tyralis et al., 1 Mar 2026).

Strict consistency follows from the population risk

yˉi:=Ti1t=1Tiyi,t,\bar y_i := T_i^{-1} \sum_{t=1}^{T_i} y_{i,t},8

For each coordinate,

yˉi:=Ti1t=1Tiyi,t,\bar y_i := T_i^{-1} \sum_{t=1}^{T_i} y_{i,t},9

which vanishes if and only if

SSEi=t=1Ti(yi,ty^i,t)2,TSSi=t=1Ti(yi,tyˉi)2.\text{SSE}_i = \sum_{t=1}^{T_i} (y_{i,t} - \hat y_{i,t})^2,\qquad \text{TSS}_i = \sum_{t=1}^{T_i} (y_{i,t} - \bar y_i)^2.0

The Hessian is diagonal with entries SSEi=t=1Ti(yi,ty^i,t)2,TSSi=t=1Ti(yi,tyˉi)2.\text{SSE}_i = \sum_{t=1}^{T_i} (y_{i,t} - \hat y_{i,t})^2,\qquad \text{TSS}_i = \sum_{t=1}^{T_i} (y_{i,t} - \bar y_i)^2.1, so the risk is strictly convex in SSEi=t=1Ti(yi,ty^i,t)2,TSSi=t=1Ti(yi,tyˉi)2.\text{SSE}_i = \sum_{t=1}^{T_i} (y_{i,t} - \hat y_{i,t})^2,\qquad \text{TSS}_i = \sum_{t=1}^{T_i} (y_{i,t} - \bar y_i)^2.2 and the minimizer is unique. In that sense, SSEi=t=1Ti(yi,ty^i,t)2,TSSi=t=1Ti(yi,tyˉi)2.\text{SSE}_i = \sum_{t=1}^{T_i} (y_{i,t} - \hat y_{i,t})^2,\qquad \text{TSS}_i = \sum_{t=1}^{T_i} (y_{i,t} - \bar y_i)^2.3 strictly elicits SSEi=t=1Ti(yi,ty^i,t)2,TSSi=t=1Ti(yi,tyˉi)2.\text{SSE}_i = \sum_{t=1}^{T_i} (y_{i,t} - \hat y_{i,t})^2,\qquad \text{TSS}_i = \sum_{t=1}^{T_i} (y_{i,t} - \bar y_i)^2.4 (Tyralis et al., 1 Mar 2026).

The functional is also strictly identifiable with identification function

SSEi=t=1Ti(yi,ty^i,t)2,TSSi=t=1Ti(yi,tyˉi)2.\text{SSE}_i = \sum_{t=1}^{T_i} (y_{i,t} - \hat y_{i,t})^2,\qquad \text{TSS}_i = \sum_{t=1}^{T_i} (y_{i,t} - \bar y_i)^2.5

Its empirical analog averages these weighted residual vectors either column-wise or row-wise, depending on the orientation of the data matrix. This places the Nash–Sutcliffe functional within the standard joint framework of strictly consistent scoring and strict identification (Tyralis et al., 1 Mar 2026).

4. Relation to average NSE and implications for estimation

For a single series, as SSEi=t=1Ti(yi,ty^i,t)2,TSSi=t=1Ti(yi,tyˉi)2.\text{SSE}_i = \sum_{t=1}^{T_i} (y_{i,t} - \hat y_{i,t})^2,\qquad \text{TSS}_i = \sum_{t=1}^{T_i} (y_{i,t} - \bar y_i)^2.6, SSEi=t=1Ti(yi,ty^i,t)2,TSSi=t=1Ti(yi,tyˉi)2.\text{SSE}_i = \sum_{t=1}^{T_i} (y_{i,t} - \hat y_{i,t})^2,\qquad \text{TSS}_i = \sum_{t=1}^{T_i} (y_{i,t} - \bar y_i)^2.7 consistently estimates SSEi=t=1Ti(yi,ty^i,t)2,TSSi=t=1Ti(yi,tyˉi)2.\text{SSE}_i = \sum_{t=1}^{T_i} (y_{i,t} - \hat y_{i,t})^2,\qquad \text{TSS}_i = \sum_{t=1}^{T_i} (y_{i,t} - \bar y_i)^2.8. The population counterpart of SSEi=t=1Ti(yi,ty^i,t)2,TSSi=t=1Ti(yi,tyˉi)2.\text{SSE}_i = \sum_{t=1}^{T_i} (y_{i,t} - \hat y_{i,t})^2,\qquad \text{TSS}_i = \sum_{t=1}^{T_i} (y_{i,t} - \bar y_i)^2.9 becomes

NSEi=1SSEiTSSi,LNS,i=1NSEi=SSEiTSSi.\text{NSE}_i = 1 - \frac{\text{SSE}_i}{\text{TSS}_i},\qquad L_{\text{NS},i} = 1 - \text{NSE}_i = \frac{\text{SSE}_i}{\text{TSS}_i}.0

Maximizing NSE is therefore equivalent to minimizing expected scaled squared error. If NSEi=1SSEiTSSi,LNS,i=1NSEi=SSEiTSSi.\text{NSE}_i = 1 - \frac{\text{SSE}_i}{\text{TSS}_i},\qquad L_{\text{NS},i} = 1 - \text{NSE}_i = \frac{\text{SSE}_i}{\text{TSS}_i}.1 is a function of covariates NSEi=1SSEiTSSi,LNS,i=1NSEi=SSEiTSSi.\text{NSE}_i = 1 - \frac{\text{SSE}_i}{\text{TSS}_i},\qquad L_{\text{NS},i} = 1 - \text{NSE}_i = \frac{\text{SSE}_i}{\text{TSS}_i}.2, the minimizer is the conditional mean NSEi=1SSEiTSSi,LNS,i=1NSEi=SSEiTSSi.\text{NSE}_i = 1 - \frac{\text{SSE}_i}{\text{TSS}_i},\qquad L_{\text{NS},i} = 1 - \text{NSE}_i = \frac{\text{SSE}_i}{\text{TSS}_i}.3, because the scale factor does not change the argmin (Tyralis et al., 1 Mar 2026).

For multiple series,

NSEi=1SSEiTSSi,LNS,i=1NSEi=SSEiTSSi.\text{NSE}_i = 1 - \frac{\text{SSE}_i}{\text{TSS}_i},\qquad L_{\text{NS},i} = 1 - \text{NSE}_i = \frac{\text{SSE}_i}{\text{TSS}_i}.4

which is the population analog of NSEi=1SSEiTSSi,LNS,i=1NSEi=SSEiTSSi.\text{NSE}_i = 1 - \frac{\text{SSE}_i}{\text{TSS}_i},\qquad L_{\text{NS},i} = 1 - \text{NSE}_i = \frac{\text{SSE}_i}{\text{TSS}_i}.5. Hence the common practice of maximizing average NSE is the sample analog of minimizing expected Nash–Sutcliffe loss. At the same time, the paper emphasizes the implicit assumption behind simply averaging NSEi=1SSEiTSSi,LNS,i=1NSEi=SSEiTSSi.\text{NSE}_i = 1 - \frac{\text{SSE}_i}{\text{TSS}_i},\qquad L_{\text{NS},i} = 1 - \text{NSE}_i = \frac{\text{SSE}_i}{\text{TSS}_i}.6: in the NSEi=1SSEiTSSi,LNS,i=1NSEi=SSEiTSSi.\text{NSE}_i = 1 - \frac{\text{SSE}_i}{\text{TSS}_i},\qquad L_{\text{NS},i} = 1 - \text{NSE}_i = \frac{\text{SSE}_i}{\text{TSS}_i}.7 orientation, it treats the NSEi=1SSEiTSSi,LNS,i=1NSEi=SSEiTSSi.\text{NSE}_i = 1 - \frac{\text{SSE}_i}{\text{TSS}_i},\qquad L_{\text{NS},i} = 1 - \text{NSE}_i = \frac{\text{SSE}_i}{\text{TSS}_i}.8 series as realizations of a single non-stationary stochastic process (Tyralis et al., 1 Mar 2026).

This observation is central to the status of the Nash–Sutcliffe functional. The functional is not merely a reformulation of a familiar score; it identifies the exact population target being optimized when average NSE is used as an evaluation criterion. A plausible implication is that disagreements between training under MSE and evaluation under average NSE are not incidental but target mismatch: the two procedures optimize different functionals unless the weighting induced by NSEi=1SSEiTSSi,LNS,i=1NSEi=SSEiTSSi.\text{NSE}_i = 1 - \frac{\text{SSE}_i}{\text{TSS}_i},\qquad L_{\text{NS},i} = 1 - \text{NSE}_i = \frac{\text{SSE}_i}{\text{TSS}_i}.9 is immaterial.

5. Nash–Sutcliffe regression and matrix orientation

For series NSEi\text{NSE}_i0 with design NSEi\text{NSE}_i1 and response NSEi\text{NSE}_i2, Nash–Sutcliffe linear regression minimizes

NSEi\text{NSE}_i3

This is weighted least squares with series-specific weights NSEi\text{NSE}_i4. The normal equations are

NSEi\text{NSE}_i5

with solution

NSEi\text{NSE}_i6

The paper presents this as the multi-series analog of the closed-form Nash–Sutcliffe regression estimators, and notes that serial dependence within series suggests robust variance estimation for inference, such as HAC estimators, if parametric uncertainty is needed (Tyralis et al., 1 Mar 2026).

Two matrix orientations are distinguished. In the NSEi\text{NSE}_i7 setting, each column is one NSEi\text{NSE}_i8-dimensional realization, such as a fixed-length time series, of a single NSEi\text{NSE}_i9-dimensional random vector yˉi\bar y_i0. In the yˉi\bar y_i1 setting, each row is one yˉi\bar y_i2-dimensional realization at a single time, allowing multiple stationary, dependent time series with differing properties across series. The paper proposes the yˉi\bar y_i3 reorientation for forecasting and states that it is a more natural empirical implementation of the NSE than the earlier formulation (Tyralis et al., 1 Mar 2026).

Under the yˉi\bar y_i4 orientation, the realized Euclidean loss is

yˉi\bar y_i5

and the realized Nash–Sutcliffe loss is

yˉi\bar y_i6

Section 5.6 shows that minimizing realized yˉi\bar y_i7 across columns with a shared linear model decouples into yˉi\bar y_i8 separate least-squares problems in this orientation; when weights are constant with respect to each column’s parameters, the argmin matches that of multivariate least-squares for each column. In large datasets, this alignment supports global models that share parameters across series, and the reported simulations and hydrometeorological applications show that Nash–Sutcliffe regression can yield substantially lower realized Nash–Sutcliffe losses than one-dimensional local regressions and multivariate OLS, especially for streamflow (Tyralis et al., 1 Mar 2026).

6. Assumptions, extensions, and interpretive cautions

The population theory requires finite component-wise second moments, yˉi\bar y_i9, and finite LNS,iL_{\text{NS},i}0. The denominator must be strictly positive almost surely, so the support excludes

LNS,iL_{\text{NS},i}1

In samples, LNS,iL_{\text{NS},i}2, equivalently LNS,iL_{\text{NS},i}3, is required for each series. Degenerate series that are constant over time make NSE undefined and LNS,iL_{\text{NS},i}4 infinite. To handle near-zero denominators, the paper introduces the extended Nash–Sutcliffe loss

LNS,iL_{\text{NS},i}5

which elicits

LNS,iL_{\text{NS},i}6

This ensures positivity of the denominator while slightly changing the target functional (Tyralis et al., 1 Mar 2026).

A separate hydrological critique is directed at NSE itself rather than at the decision-theoretic reconstruction. Using the Murphy (1988) and Gupta et al. (2009) decomposition,

LNS,iL_{\text{NS},i}7

so that NSE is an entangled function of bias, variability mismatch, and correlation. In controlled experiments, NSE is least sensitive to bias, more sensitive to variability mismatch, and most sensitive to correlation loss; the degradation is strongly nonlinear for bias and variability, but approximately linear for correlation errors. One consequence is that the same NSE value can correspond to very different error realities, making threshold heuristics unreliable. Within the cited study’s framework, these properties can alter the model solution space, sampling sufficiency status, and inferred runoff-generation hypotheses, which is why that study recommends KGEss rather than NSE as a single-metric choice for model evaluation and hypothesis testing (Khatami et al., 2020).

The Nash–Sutcliffe functional does not remove those interpretive cautions about the score. Instead, it clarifies what is being optimized when average NSE is used. The practical guidance given in the decision-theoretic treatment is correspondingly narrow and operational: treat NSE as a skill score relative to the mean; average NSE only across series that plausibly share an underlying stochastic structure, or use the LNS,iL_{\text{NS},i}8 reorientation for forecasting; if evaluation is by average NSE, train with Nash–Sutcliffe loss; and use LNS,iL_{\text{NS},i}9 to mitigate near-zero denominators (Tyralis et al., 1 Mar 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Nash-Sutcliffe Functional.