Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sparse Isotonic Shapley Regression (SISR)

Updated 4 December 2025
  • SISR is a framework that restores additivity in Shapley values through a learned monotonic transformation of coalition payoffs.
  • It alternates between isotonic regression and normalized hard-thresholding to enforce exact ℓ0 sparsity in attributions.
  • Empirical evaluations show that SISR delivers stable and interpretable explanations even with heavy-tailed, dependent features.

Sparse Isotonic Shapley Regression (SISR) addresses two foundational obstacles in Shapley-based model explainability: the violation of the canonical additivity assumption in real-world payoff constructions, and the need for sparse, interpretable attributions in high-dimensional feature spaces. The method provides a unified framework that simultaneously estimates a monotonic, data-driven transformation to restore additivity of the worth function and enforces an ℓ0\ell_0 sparsity constraint on the Shapley vector, thereby yielding more consistent and computationally efficient explanations in settings plagued by non-Gaussianity, heavy tails, feature dependence, or idiosyncratic loss scales (She, 2 Dec 2025).

1. Model Formulation and Objectives

SISR considers a family of coalition values {νA:A⊆{1,…,p}}\{\nu_A : A\subseteq\{1,\ldots,p\}\} and seeks to jointly: (i) learn a strictly increasing scalar function f:R→Rf:\mathbb{R}\to\mathbb{R} with f(0)=0f(0)=0, such that f(νA)f(\nu_A) becomes additive in the underlying Shapley scores; (ii) enforce sparsity (ℓ0\ell_0 constraint) on the Shapley vector φ\varphi, aligning with the hypothesis that only a small subset of features is truly relevant.

The model assumes that for all coalitions AA,

f(νA)≈∑j∈Aφj∗+noiseA,noiseA∼N(0,σA2)f(\nu_A) \approx \sum_{j\in A} \varphi_j^* + \text{noise}_A, \quad \text{noise}_A \sim N(0,\sigma_A^2)

Enforcing ∥φ∥0≤s\|\varphi\|_0\le s directly achieves exact sparsity in attributions, bypassing the shrinkage-bias inherent to {νA:A⊆{1,…,p}}\{\nu_A : A\subseteq\{1,\ldots,p\}\}0 penalization. The primary SISR objective is:

{νA:A⊆{1,…,p}}\{\nu_A : A\subseteq\{1,\ldots,p\}\}1

Reparameterization with {νA:A⊆{1,…,p}}\{\nu_A : A\subseteq\{1,\ldots,p\}\}2 and {νA:A⊆{1,…,p}}\{\nu_A : A\subseteq\{1,\ldots,p\}\}3 (discrete), and introducing matrix {νA:A⊆{1,…,p}}\{\nu_A : A\subseteq\{1,\ldots,p\}\}4 (rows index coalitions), yields the equivalent quadratic problem with order restrictions on {νA:A⊆{1,…,p}}\{\nu_A : A\subseteq\{1,\ldots,p\}\}5.

2. Optimization Algorithm

SISR employs alternating minimization over {νA:A⊆{1,…,p}}\{\nu_A : A\subseteq\{1,\ldots,p\}\}6 (sparse Shapley vector) and {νA:A⊆{1,…,p}}\{\nu_A : A\subseteq\{1,\ldots,p\}\}7 (discretized monotonic transform):

A. Monotonic Transformation via Isotonic Regression:

With {νA:A⊆{1,…,p}}\{\nu_A : A\subseteq\{1,\ldots,p\}\}8 fixed, update {νA:A⊆{1,…,p}}\{\nu_A : A\subseteq\{1,\ldots,p\}\}9 by solving a weighted isotonic regression problem:

f:R→Rf:\mathbb{R}\to\mathbb{R}0

This is executed exactly using the Pool-Adjacent-Violators (PAV) algorithm with computational complexity f:R→Rf:\mathbb{R}\to\mathbb{R}1, for f:R→Rf:\mathbb{R}\to\mathbb{R}2 (often reduced via subsampling). The output vector f:R→Rf:\mathbb{R}\to\mathbb{R}3 provides the transformed payoffs at observed coalition values.

B. Sparsity via Normalized Hard-Thresholding:

With f:R→Rf:\mathbb{R}\to\mathbb{R}4 fixed, update f:R→Rf:\mathbb{R}\to\mathbb{R}5 by minimizing the objective under f:R→Rf:\mathbb{R}\to\mathbb{R}6 constraints. The step utilizes a quadratic surrogate f:R→Rf:\mathbb{R}\to\mathbb{R}7 and a hard-thresholding operator f:R→Rf:\mathbb{R}\to\mathbb{R}8, where f:R→Rf:\mathbb{R}\to\mathbb{R}9 zeros all but the f(0)=0f(0)=00 largest f(0)=0f(0)=01. The global minimizer takes the form:

f(0)=0f(0)=02

where f(0)=0f(0)=03, f(0)=0f(0)=04.

Algorithm Sketch:

  1. Initialize f(0)=0f(0)=05 as unit-norm f(0)=0f(0)=06-sparse vector, f(0)=0f(0)=07; precompute f(0)=0f(0)=08.
  2. Alternate: (a) update f(0)=0f(0)=09 by normalized hard-thresholding; (b) update f(νA)f(\nu_A)0 with weighted PAVA.
  3. Continue until objective stabilizes.

Convergence properties are established: each f(νA)f(\nu_A)1-update reduces objective, the sequence converges, and gaps vanish if f(νA)f(\nu_A)2 is strictly greater than f(νA)f(\nu_A)3.

3. Theoretical Guarantees and Analysis

Under the generative model f(νA)f(\nu_A)4, with f(νA)f(\nu_A)5-sparsity of f(νA)f(\nu_A)6 and f(νA)f(\nu_A)7, the following results are obtained:

  • Consistency: As noise vanishes or sample size of coalition values increases, the isotonic regression step recovers the true monotonic transformation f(νA)f(\nu_A)8 up to negligible error.
  • Support Recovery: Provided the minimum nonzero f(νA)f(\nu_A)9 is substantially larger than the noise and the signal-to-noise ratio ℓ0\ell_00 is large, the normalized hard-thresholding step identifies the correct support with high probability.
  • Feature Dependence and Heavy-Tailed Payoffs: SISR learns a transform ℓ0\ell_01 that absorbs the nonlinearities, restoring additivity even for payoff structures arising from feature dependence, heavy tails, or domain-specific loss constructions (e.g., ℓ0\ell_02 under correlated regressors). The canonical linear-Gaussian assumption underlying standard Shapley can otherwise suffer severe violations leading to distorted attributions.

4. Empirical Performance

Comprehensive experiments on synthetic and real data characterize SISR performance:

  • Synthetic Recovery: For a battery of toy monotonic transforms (e.g., square root, fifth root, logarithm, exponential, tanh-like, Gaussian CDF-odds), SISR accurately fits ℓ0\ell_03 (correlation ℓ0\ell_04).
  • Sparse Support Recovery: For synthetic ℓ0\ell_05 with ℓ0\ell_06, ℓ0\ell_07 up to 25, and noise up to ℓ0\ell_08, both support agreement and affinity ℓ0\ell_09 remain near 1. Recovery degrades gracefully with noise; e.g., for φ\varphi0, true support is recovered approximately 90% of the time.
  • φ\varphi1 Payoff Scenarios: Simulations with φ\varphi2 payoff under feature correlation/irrelevance demonstrate that raw φ\varphi3 is not additive; SISR fits a correction, delivering stable attributions, while standard Shapley responses become highly nonrobust.
  • Real Data Applications:
    • Prostate dataset: Shapley values exaggerate the importance of “svi”; SISR zeros it out, aligning with AIC/BIC and expert knowledge.
    • Boston housing: With robust payoff φ\varphi4, baseline Shapley values flip important variable signs; SISR’s log-like transform yields stable rankings.
    • Credit and Pima diabetes: Across diverse payoff schemes, SISR provides consistent sparse attributions (selected via RIC), whereas standard SAGE/Shapley attributions are unstable.

5. Computational and Practical Considerations

  • Complexity: Weighted PAVA is φ\varphi5 per iteration. When utilizing sparse storage for φ\varphi6 (each coalition row φ\varphi7 nonzeros), φ\varphi8-updates cost φ\varphi9 per iteration. Convergence is typically reached within tens of alternations.
  • Sparsity Level Selection: If AA0 is unknown, Risk Inflation Criterion (RIC) can be used to choose AA1. Empirical evidence indicates reliable selection.
  • Initialization: AA2 as the top-AA3 entries of AA4 (normalized) and AA5 (with large AA6 for numerical stability) are recommended.
  • Infinite Weights for Special Coalitions: To ensure AA7 and the “efficiency” constraint, theoretical Shapley weights AA8 and AA9 diverge; in practice, substituting large multiples (e.g., f(νA)≈∑j∈Aφj∗+noiseA,noiseA∼N(0,σA2)f(\nu_A) \approx \sum_{j\in A} \varphi_j^* + \text{noise}_A, \quad \text{noise}_A \sim N(0,\sigma_A^2)010 or f(νA)≈∑j∈Aφj∗+noiseA,noiseA∼N(0,σA2)f(\nu_A) \approx \sum_{j\in A} \varphi_j^* + \text{noise}_A, \quad \text{noise}_A \sim N(0,\sigma_A^2)1100) of the maximum weight is effective.
  • Extensibility: Alternating PAVA plus hard-thresholding can extend to other convex loss functions (e.g., GLMs with canonical links) by substituting the quadratic surrogate with a suitable local approximation.

6. Significance and Implications in Explainable AI

SISR advances the Shapley explanation paradigm by jointly resolving two stated limitations: First, it restores additivity in the presence of nonadditive, domain-specific, or statistically problematic payoff functions via a learned monotonic transformation. Second, it provides exact f(νA)≈∑j∈Aφj∗+noiseA,noiseA∼N(0,σA2)f(\nu_A) \approx \sum_{j\in A} \varphi_j^* + \text{noise}_A, \quad \text{noise}_A \sim N(0,\sigma_A^2)2-sparse attributions in high-dimensional settings, overcoming the inconsistency and computational burden of post hoc thresholding.

Empirical evidence demonstrates that SISR explanations are both more stable and faithful to domain structure—eliminating attributions to irrelevant features and correcting severe rank/sign distortions observed in conventional Shapley or SAGE values under nonlinear or unstable payoffs.

A plausible implication is that SISR is well-positioned as a general-purpose attribution framework for scenarios involving heavy tails, feature dependence, or overarching nonlinearity in payoff structure, where standard linear Shapley values are provably unreliable (She, 2 Dec 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Sparse Isotonic Shapley Regression (SISR).