---
title: 'Shapley Share: Unifying Allocation in Games'
url: https://www.emergentmind.com/topics/shapley-share
type: topic
---

# Shapley Share: Unifying Allocation in Games

Searching arXiv for recent and foundational papers on "Shapley share" and related Shapley-value allocation settings.
In the literature surveyed here, **Shapley Share** denotes a Shapley-value-based allocation of a coalition-generated quantity to individual entities. The quantity being allocated varies by domain—product profit, route cost, network coverage, predictive utility, explained variance, identification power, or instancewise model output utility—but the common mechanism is the same: each player receives the average marginal contribution it makes over coalition formation orders. This makes “share” a unifying interpretation rather than a single standardized object: in some papers it is a fair share of profit or influence, in others a fair share of cost, and in others a local attribution score in the same units as the underlying utility [2606.24121][1909.04713][2001.09593][1808.02610].

## 1. Cooperative-game foundation

The canonical formalization is a cooperative game \((N,v)\), where \(N\) is the player set and \(v:2^N\to\mathbb R\) is the characteristic function. The Shapley value of player \(i\) is the weighted average of marginal contributions
\[
\phi_i(v)=\sum_{S\subseteq N\setminus\{i\}} \frac{|S|!(|N|-|S|-1)!}{|N|!}\bigl(v(S\cup\{i\})-v(S)\bigr).
\]
Equivalent permutation forms interpret \(\phi_i(v)\) as the expected increase in value caused by inserting \(i\) into a random predecessor coalition [2606.24121][2306.02071].

Under this definition, a “Shapley share” is additive in the units of the underlying game. In network coverage, it is a node’s fair share of total covered vertices; in collaborative machine learning, it is a data owner’s fair share of model utility; in ride-sharing and congestion games, it is a player’s share of total cost rather than surplus [2606.24121][2306.02071][1909.04713][1710.01634]. The standard axiomatic basis remains the familiar one: efficiency, symmetry, dummy or null-player behavior, and additivity or linearity. Efficiency is especially central because it makes the shares exhaustive:
\[
\sum_{i\in N}\phi_i(v)=v(N)-v(\emptyset)
\]
or, in pure cost-sharing settings, the corresponding total cost identity [1808.02610][1710.01634].

This general scheme does not require the share to be normalized to \([0,1]\). Several papers explicitly use the Shapley value as an additive attribution score in the same units as coverage, utility, or cost, with normalization only as an optional post-processing step [2606.24121][1909.04713].

## 2. Principal semantic regimes

Across application areas, the same Shapley mechanism is attached to different kinds of players and different characteristic functions. This suggests that “Shapley Share” is best understood as a transferable allocation principle rather than a domain-specific formula.

| Domain | Players | Shared quantity |
|---|---|---|
| Structured model interpretation | Features | Instancewise predictive utility gain |
| Dataset valuation | Datasets or data points | ML task utility |
| Social/network analysis | Nodes or seeds | Coverage or expected influence |
| Cost sharing | Passengers or resource users | Route cost or resource cost |
| Statistical attribution | Covariates or variables | \(R^2\) or identification power |

In structured feature attribution, the game value is an instancewise predictive utility. “L-Shapley and C-Shapley” define
\[
v_x(S)
\]
from predictive log-probability and distribute the utility gain \(v_x([d])-v_x(\emptyset)\) across features. The share is therefore a local attribution of model output, not a global population statistic [1808.02610]. In classification data valuation, “CS-Shapley” changes the characteristic function itself to
\[
v_{y_i}(S)=a_S(D_{y_i})e^{a_S(D_{-y_i})},
\]
so that a datum’s share reflects class-specific help or harm rather than only global validation accuracy [2211.06800].

In dataset valuation more broadly, the players are data owners and the coalition value is the performance of the model trained on pooled data. “DU-Shapley” preserves the standard semantic target—average marginal contribution share—but replaces combinatorial averaging with a size-driven proxy under \(u(\mathcal S)=w(n_{\mathcal S})\) [2306.02071]. “An Asymptotic Analysis of the Shapley Value for Dataset Valuation” studies the same notion in a large-player regime and shows that a fixed dataset’s Shapley share is asymptotically governed by dataset size, a first-order utility signal, and a harmonic scaling factor of order \(\log I / I\) [2607.03374].

In network analysis, the shared quantity is often coverage or influence. “Sphere of Influence Centrality via Shapley Values” defines a coalition’s value as the size of the covered set under single-hop, \(k\)-hop, or multi-path connectivity, so that each node’s Shapley value is its fair share of total network coverage [2606.24121]. “Efficient Shapley-Based Influence Attribution in Social Networks” fixes a seed set \(T\) under the Independent Cascade model and allocates the expected number of activated non-seed nodes among those seeds, yielding an explicitly ex-ante share of influence [2605.28086].

In cost-sharing, the sign of the game is reversed but the semantics are parallel. “Fair Sharing: The Shapley Value for Ride-Sharing and Routing Games” models coalition value as route cost and uses the Shapley value to split the common ride cost among passengers [1909.04713]. “Computing Approximate Pure Nash Equilibria in Shapley Value Weighted Congestion Games” defines the atomic Shapley share \(\chi_e(i,A)\) as player \(i\)’s expected marginal increase in the cost of resource \(e\) when the users are \(A\) [1710.01634].

In statistical attribution, the shared quantity can be explained variance or identification power. “Shapley value confidence intervals for attributing variance explained” decomposes model \(R^2\) into covariate-level shares [2001.09593]. “What makes you unique?” defines a subject-specific value
\[
val(u)=\log_2\frac{n}{N_t(u)}
\]
and allocates the reduction in log-cardinality of the candidate set across variables, so that each variable receives a Shapley share of identification power [2105.08013].

## 3. Structural extensions and alternative notions of share

A major line of work generalizes the object being shared. Flores, Molina, and Tejada define the **Shapley group value**
\[
\phi^g(C;N,v)=\phi_{\mathbf c}(N_C,v_C),
\]
where coalition \(C\) is merged into a proxy player \(\mathbf c\). This is a group-level share: not the sum of members’ individual shares, but the Shapley value of the group when acting as one unit [1412.5429]. The distinction matters because group valuation can capture redundancy and complementarity that vanish under simple summation.

A still stronger extension appears in the Hodge-theoretic literature. “A Hodge theoretic extension of Shapley axioms” assigns each player a full component game \(v_i(S)\), so that for every coalition \(S\),
\[
v(S)=\sum_i v_i(S),
\]
and \(v_i([N])\) recovers the classical Shapley value. In this formulation, a Shapley share is not only a scalar payoff at the grand coalition but a coalition-wise fair allocation profile across the entire lattice of coalitions [2106.15094].

Other papers alter the meaning of marginal contribution itself. “A Ratio-Based Shapley Value for Collaborative Machine Learning” replaces additive increments by relative gains,
\[
\Delta^{rel}_{i,C}=\frac{v_{C\cup\{i\}}}{v_C}-1
\]
when \(v_C\neq 0\), producing a proportional rather than additive notion of share. The resulting reward rule keeps the incentive framework of earlier collaborative-ML work but changes the normative meaning of contribution from absolute improvement to relative uplift [2510.13261]. “Absolute Shapley Value” changes semantics in a different direction by replacing signed marginal contributions with their absolute values; the paper’s own interpretation implies that this is better read as an importance-magnitude score than as a canonical compensation share [2003.10076].

The power-index literature overlaps with this vocabulary. “The Spread of Voting Attitudes in Social Networks” begins from the Shapley–Shubik power index
\[
p(I)=\pi_I/n!
\]
and then defines a graph-local power rule in which winning-side voters in a neighborhood receive reciprocal local power. This is not a classical cooperative-game Shapley value, but it is a clear Shapley-style distribution of voting power across agents [1812.02143].

## 4. Computation, locality, and scalability

Exact Shapley computation is combinatorial. The recurring computational theme is therefore to exploit structure—graph locality, restricted supports, coalition-size surrogates, dynamic reuse, or closed forms—to make shares tractable.

For graph-structured feature attribution, “L-Shapley and C-Shapley” replace global coalition averaging by local or connected coalition averaging. On a line graph, all \(d\) L-Shapley values require about
\[
2^{2k}d
\]
model evaluations, while C-Shapley on a line graph requires
\[
\mathcal O(k^2 d)
\]
evaluations by enumerating only connected intervals; on a grid, L-Shapley grows as \(2^{4k^2}d\) [1808.02610]. In dataset valuation, DU-Shapley reduces the support of the expectation from \(2^I\) coalition terms to only \(I\) regularly spaced aggregate-size states [2306.02071].

In dynamic settings, “Dynamic Shapley Computation” represents valuation as a player-by-task matrix \(\Phi_{i,j}=\phi(z_i,t_j)\) and updates that matrix locally. The paper reports task updates in milliseconds and player updates cheaper by up to three orders of magnitude than full recomputation, based on utility locality and coalition locality [2605.20620].

Network influence papers exhibit a sharp tractability boundary. In the single-step Independent Cascade setting, exact Shapley values for all seeds can be computed in
\[
O(|V|\cdot |T|^3),
\]
whereas for \(K\)-step termination with \(K\ge 2\) the exact computation becomes \(\#P\)-hard; additive-error and RR-set-based approximations are then used instead [2605.28086]. For sphere-of-influence coverage games, closed-form polynomial-time formulas are available for several reachability rules, and empirical approximation ratios for top-\(m\) Shapley ranking approach \(0.9\) in the reported experiments [2606.24121].

Patent valuation pushes locality into a graph-conditioned regime. “A Framework for Graph-Conditioned Hierarchical Shapley Attribution in Patent Valuation” restricts each patent’s coalition to its Markov Blanket in a knowledge graph, motivated by the C-SVE conditional-independence theorem. At \(n=100\), the paper reports median blanket size of 32.9 percent of \(n\), 90th-percentile blanket size of 55.2 percent of \(n\), and runtime of 10 milliseconds per patent; the difference against exact ground truth at \(n=12\) is 0.088, and the difference against a high-sample Monte Carlo reference at \(n=100\) is \(0.062 \pm 0.003\) [2606.01632].

Cost-sharing problems also separate structurally easy and hard cases. In ride-sharing with a fixed drop-off order, exact Shapley-based cost shares are computable efficiently; when every coalition is allowed to choose its own shortest route, exact computation is not efficiently computable unless \(P=NP\), and the paper advocates a fixed-order proxy called SHAPO [1909.04713].

## 5. Statistical, asymptotic, and uncertainty-aware formulations

A mature treatment of Shapley shares requires more than point estimation. Several papers therefore study sampling variability, asymptotic scale, and stochastic contributors.

For regression attribution, “Shapley value confidence intervals for attributing variance explained” proves asymptotic normality of Shapley values for \(R^2\) decomposition under a pseudo-elliptical assumption and develops plug-in confidence intervals and pairwise hypothesis tests [2001.09593]. This turns a descriptive share of explained variance into an inferentially testable quantity.

For stochastic data providers, “Shapley Value on Uncertain Data” treats each player’s Shapley value itself as a random variable induced by sampling from an underlying distribution. The primary objects become
\[
\mathbb E[\phi_i] \quad \text{and} \quad \mathrm{Var}(\phi_i),
\]
with unbiased estimators for both. The paper proposes baseline, pooled, and stratified pooled Monte Carlo schemes, and reports substantial variance reduction from the stratified pooled variant at minimal additional cost [2601.14543]. In this setting, a Shapley share is no longer merely a point allocation but an expectation-plus-uncertainty object.

Large-population asymptotics for dataset valuation sharpen this perspective. “An Asymptotic Analysis of the Shapley Value for Dataset Valuation” shows that the exact Shapley share of a fixed dataset is asymptotically captured by a leading term involving dataset size, a first-order utility signal, and the harmonic factor \(H_{I-1}/I\), implying scale \(\log I/I\) for nondegenerate contributors [2607.03374]. This places approximate methods such as DU-Shapley into a common asymptotic frame and clarifies what relative accuracy should mean when individual shares vanish with market size.

## 6. Caveats, boundaries, and contested interpretations

The phrase **Shapley Share** is not standardized across fields. Related papers often speak instead of the Shapley value, Shapley group value, Shapley–Shubik power, cost shares, or local attribution scores [1412.5429][1812.02143][1710.01634]. This suggests that the phrase is best used as an umbrella descriptor for Shapley-based allocation rather than as the name of a single canonical construct.

The share is only as meaningful as the characteristic function \(v(S)\). In feature attribution, the value depends on how missing features are operationalized; the plug-in masking scheme used in L-Shapley and C-Shapley can generate out-of-distribution inputs, and locality or connectedness assumptions may fail if long-range interactions are strong [1808.02610]. In network centrality and influence attribution, a Shapley ranking is an individual score, whereas the actual optimization target is subset coverage; high-share nodes can still overlap heavily, so top-\(m\) Shapley selection is only a heuristic with guarantees, not the exact combinatorial optimum [2606.24121].

Once exogenous corrections or alternative marginal semantics are introduced, classical axiomatic interpretation becomes more fragile. The ratio-based collaborative-ML proposal preserves the cited incentive conditions but changes the value function itself, so it offers a different notion of contribution rather than a re-expression of the additive Shapley rule [2510.13261]. The improved Shapley allocation for the we-media value chain layers AHP-derived innovation weights onto a Shapley-style formula; in the worked example the reported allocations exceed the total coalition profit, so the resulting object is best regarded as a Shapley-inspired adjusted allocation heuristic rather than a pure Shapley share in the classical efficient sense [2412.18130].

These caveats do not weaken the central importance of the concept. Rather, they delimit its scope. A Shapley share is most coherent when the coalition value is well specified, the interpretation of players and coalitions is stable, and the computational approximation respects the structure of the game being modeled. Under those conditions, it remains one of the most portable and technically disciplined ways to allocate interaction-generated value across agents, features, datasets, costs, or coalitions [2306.02071][1808.02610][1909.04713].

Source: https://www.emergentmind.com/topics/shapley-share