Papers
Topics
Authors
Recent
Search
2000 character limit reached

Semiparametric Inference for Conditional Shapley Feature Importance

Published 9 Sep 2026 in stat.ME | (2609.10313v1)

Abstract: Shapley values are widely used for post-hoc feature attribution, but most estimators return point quantities and do not quantify uncertainty, and popular implementations sample out-of-coalition features from their marginal distribution, which misattributes importance when features are dependent. This paper studies the conditional formulation, in which out-of-coalition features are integrated out under their true conditional distribution. The target is a global, loss-based importance that pairs a conditional value function with a SAGE-style loss aggregation. We propose a one-step estimator with K-fold cross-fitting and a U-statistic correction of the squared loss that removes the Monte Carlo bias of the naive plug-in; it is n\sqrt{n}-consistent and asymptotically normal under double-robust rate conditions, and the resulting Wald interval attains nominal coverage. A Pinsker-type bound quantifies the bias from misspecifying the working copula class, while vine copulas keep conditional sampling tractable. In a Gaussian design study with n= 500, the empirical coverage of the 95% interval lies between 0.91 and 0.96 across all features, the test holds its Type-I rate at 0.05, and it reaches power one for moderate signals. Applied to the UCI Concrete and California Housing data, the method identifies the conditionally informative features with Bonferroni-controlled significance.

Authors (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.