Nash–Sutcliffe Functional and Loss Framework
- The Nash–Sutcliffe functional is a data-weighted component-wise mean serving as an elicitable target, linking hydrological scoring with a decision-theoretic foundation.
- It reweights observations via the inverse centered sum-of-squares, ensuring strict consistency and identifiability when optimizing predictive performance.
- Its framework underpins methods like Nash–Sutcliffe regression, aligning average NSE evaluation with expected loss minimization across multiple series.
Searching arXiv for the cited papers and closely related NSE literature. The Nash–Sutcliffe functional is the elicitable and identifiable multi-dimensional target that is strictly elicited by the Nash–Sutcliffe loss, . In the formulation developed in "Learning with the Nash-Sutcliffe loss," it is a data-weighted component-wise mean, obtained by reweighting the underlying distribution by the inverse centered within-vector sum of squares. This construction gives a decision-theoretic foundation to NSE-based forecast evaluation and model estimation across multiple series, and clarifies the conditions under which maximizing average NSE is the sample analog of minimizing an expected loss (Tyralis et al., 1 Mar 2026).
1. From Nash–Sutcliffe efficiency to Nash–Sutcliffe loss
For a single series with observations , forecasts , and sample mean
the standard sums of squares are
The Nash–Sutcliffe efficiency and its negatively oriented counterpart are
In this form, compares predictive skill to the naïve mean baseline that always predicts , while is the corresponding relative squared-error loss. The condition 0 ensures 1 (Tyralis et al., 1 Mar 2026).
Across 2 series, the average Nash–Sutcliffe loss and average NSE are
3
In vector notation for a 4-dimensional realization 5 and prediction 6,
7
This representation is the bridge from the familiar hydrological score to a population-level scoring rule on 8 (Tyralis et al., 1 Mar 2026).
Within the hydrological literature, NSE is also characterized as
9
Its upper bound is 0, achieved only for a perfect match, while the lower bound is unbounded. NSE 1 corresponds to the naïve reference that predicts the mean 2 at all times; values 3 indicate performance worse than this baseline, and values 4 indicate improvement over the baseline (Khatami et al., 2020).
2. Decision-theoretic foundation
The decisive step is to treat the Nash–Sutcliffe loss as a scoring function in a distributional setting. Let 5 be a 6-dimensional random vector with joint distribution 7 on 8. The population loss is
9
This embeds NSE in the scoring-rules framework through a weighted Euclidean loss (Tyralis et al., 1 Mar 2026).
The underlying formal notions are those of elicitability, strict consistency, and identifiability. A functional 0 is elicitable if there exists a loss 1 such that 2 is uniquely minimized at 3. A loss is strictly 4-consistent for 5 if the expected loss is uniquely minimized at the target, and a functional is identifiable if there exists 6 such that 7 if and only if 8 (Tyralis et al., 1 Mar 2026).
The paper invokes Theorem 1 of Gneiting (2011): if 9 is a (strictly) consistent loss for an elicitable functional 0, then the weighted loss
1
is (strictly) consistent for the weighted functional
2
where 3 has density proportional to 4. Applied with 5 and the Nash–Sutcliffe weight 6, the result yields the Nash–Sutcliffe functional as the target of the Nash–Sutcliffe loss (Tyralis et al., 1 Mar 2026).
3. Definition of the Nash–Sutcliffe functional
The Nash–Sutcliffe functional is the component-wise mean under the reweighted distribution: 7 The paper explicitly describes this as a data-weighted component-wise mean (Tyralis et al., 1 Mar 2026).
Strict consistency follows from the population risk
8
For each coordinate,
9
which vanishes if and only if
0
The Hessian is diagonal with entries 1, so the risk is strictly convex in 2 and the minimizer is unique. In that sense, 3 strictly elicits 4 (Tyralis et al., 1 Mar 2026).
The functional is also strictly identifiable with identification function
5
Its empirical analog averages these weighted residual vectors either column-wise or row-wise, depending on the orientation of the data matrix. This places the Nash–Sutcliffe functional within the standard joint framework of strictly consistent scoring and strict identification (Tyralis et al., 1 Mar 2026).
4. Relation to average NSE and implications for estimation
For a single series, as 6, 7 consistently estimates 8. The population counterpart of 9 becomes
0
Maximizing NSE is therefore equivalent to minimizing expected scaled squared error. If 1 is a function of covariates 2, the minimizer is the conditional mean 3, because the scale factor does not change the argmin (Tyralis et al., 1 Mar 2026).
For multiple series,
4
which is the population analog of 5. Hence the common practice of maximizing average NSE is the sample analog of minimizing expected Nash–Sutcliffe loss. At the same time, the paper emphasizes the implicit assumption behind simply averaging 6: in the 7 orientation, it treats the 8 series as realizations of a single non-stationary stochastic process (Tyralis et al., 1 Mar 2026).
This observation is central to the status of the Nash–Sutcliffe functional. The functional is not merely a reformulation of a familiar score; it identifies the exact population target being optimized when average NSE is used as an evaluation criterion. A plausible implication is that disagreements between training under MSE and evaluation under average NSE are not incidental but target mismatch: the two procedures optimize different functionals unless the weighting induced by 9 is immaterial.
5. Nash–Sutcliffe regression and matrix orientation
For series 0 with design 1 and response 2, Nash–Sutcliffe linear regression minimizes
3
This is weighted least squares with series-specific weights 4. The normal equations are
5
with solution
6
The paper presents this as the multi-series analog of the closed-form Nash–Sutcliffe regression estimators, and notes that serial dependence within series suggests robust variance estimation for inference, such as HAC estimators, if parametric uncertainty is needed (Tyralis et al., 1 Mar 2026).
Two matrix orientations are distinguished. In the 7 setting, each column is one 8-dimensional realization, such as a fixed-length time series, of a single 9-dimensional random vector 0. In the 1 setting, each row is one 2-dimensional realization at a single time, allowing multiple stationary, dependent time series with differing properties across series. The paper proposes the 3 reorientation for forecasting and states that it is a more natural empirical implementation of the NSE than the earlier formulation (Tyralis et al., 1 Mar 2026).
Under the 4 orientation, the realized Euclidean loss is
5
and the realized Nash–Sutcliffe loss is
6
Section 5.6 shows that minimizing realized 7 across columns with a shared linear model decouples into 8 separate least-squares problems in this orientation; when weights are constant with respect to each column’s parameters, the argmin matches that of multivariate least-squares for each column. In large datasets, this alignment supports global models that share parameters across series, and the reported simulations and hydrometeorological applications show that Nash–Sutcliffe regression can yield substantially lower realized Nash–Sutcliffe losses than one-dimensional local regressions and multivariate OLS, especially for streamflow (Tyralis et al., 1 Mar 2026).
6. Assumptions, extensions, and interpretive cautions
The population theory requires finite component-wise second moments, 9, and finite 0. The denominator must be strictly positive almost surely, so the support excludes
1
In samples, 2, equivalently 3, is required for each series. Degenerate series that are constant over time make NSE undefined and 4 infinite. To handle near-zero denominators, the paper introduces the extended Nash–Sutcliffe loss
5
which elicits
6
This ensures positivity of the denominator while slightly changing the target functional (Tyralis et al., 1 Mar 2026).
A separate hydrological critique is directed at NSE itself rather than at the decision-theoretic reconstruction. Using the Murphy (1988) and Gupta et al. (2009) decomposition,
7
so that NSE is an entangled function of bias, variability mismatch, and correlation. In controlled experiments, NSE is least sensitive to bias, more sensitive to variability mismatch, and most sensitive to correlation loss; the degradation is strongly nonlinear for bias and variability, but approximately linear for correlation errors. One consequence is that the same NSE value can correspond to very different error realities, making threshold heuristics unreliable. Within the cited study’s framework, these properties can alter the model solution space, sampling sufficiency status, and inferred runoff-generation hypotheses, which is why that study recommends KGEss rather than NSE as a single-metric choice for model evaluation and hypothesis testing (Khatami et al., 2020).
The Nash–Sutcliffe functional does not remove those interpretive cautions about the score. Instead, it clarifies what is being optimized when average NSE is used. The practical guidance given in the decision-theoretic treatment is correspondingly narrow and operational: treat NSE as a skill score relative to the mean; average NSE only across series that plausibly share an underlying stochastic structure, or use the 8 reorientation for forecasting; if evaluation is by average NSE, train with Nash–Sutcliffe loss; and use 9 to mitigate near-zero denominators (Tyralis et al., 1 Mar 2026).