- The paper introduces K-Shapley values, an axiomatic merit measure for size-limited coalitions, and defines fairness as selecting arms in proportion to their estimated contributions.
- The paper’s K-SVFair-FBF algorithm combines UCB exploration, Monte Carlo permutations, repeated sampling, and randomized rounding to achieve high-probability fairness regret of approximately Õ(T^{3/4}KM^{1/2}).
- Experiments in synthetic bandits, federated learning, and influence maximization show that merit-proportional selection reduces fairness regret and can match or improve task performance versus existing baselines.
Problem setting and motivation
The paper addresses meritocratic fairness in budgeted combinatorial multi-armed bandits under full-bandit feedback (BCMAB-FBF). In this setting, a learner selects a super-arm St of size at most K from M base arms and observes only the aggregate stochastic reward Vt(St), not per-arm rewards. This contrasts with semi-bandit feedback, where individual arm rewards are visible; under full-bandit feedback, attributing reward to individual arms is fundamentally harder. The authors motivate the problem with federated learning (FL), where a server sees only global model accuracy after aggregating client updates, and social influence maximization (SIM), where only total diffusion spread is observable.
Existing fairness work in CMAB is almost entirely confined to semi-bandit feedback, achieving logarithmic regret via exposure guarantees or linear constraints. Under full-bandit feedback, only one prior work imposes fixed minimum-pull quotas per arm. The authors identify two shortcomings of that approach: quotas are typically unknown a priori, and quota-based fairness ignores merit — an arm's selection frequency should reflect its contribution rather than a static floor. Even ignoring fairness, the best known regret for bandit algorithms under full-bandit feedback is O(T2/3), so any fairness mechanism must contend with an intrinsically noisy feedback regime.
The K-Shapley value
To define merit without structural assumptions on the reward function, the paper adapts the Shapley value to budgeted coalitions. Since the valuation function V is defined only on coalitions of size at most K, classical Shapley axioms cannot be applied directly: marginal contributions to coalitions larger than K are unobservable, and the grand coalition value V([M]) is undefined. Existing extensions to restricted cooperative games do not fit — Willson-style values assign zero to infeasible coalitions, Albizuri et al.'s R-games assume grand-coalition feasibility, and graph-restricted games tie feasibility to topology rather than cardinality.
The proposed K-Shapley value K0 averages player K1's marginal contributions over all uniformly random K2-sized coalitions containing K3, weighting each by the standard Shapley combinatorial factor within that coalition. The authors prove it satisfies Symmetry, Linearity, Null Player, and a modified K-efficiency axiom requiring
K4
which reduces to classical efficiency when K5. A uniqueness theorem establishes that the K6-Shapley value is the only solution concept satisfying these four axioms, proved by extending carrier-game arguments to K7-restricted games. Meritocratic fairness is then defined as a policy whose selection probabilities are proportional to K8-Shapley values, with fairness regret measured as cumulative K9 deviation from the ideal policy M0.
Algorithms and theoretical guarantees
The paper first analyzes MURaS (Meritocratic Uniform Random Sampling), an explore-then-commit baseline that estimates M1-Shapley values via Monte Carlo permutations with repeated stochastic evaluations, then commits to the meritocratic policy. Its analysis makes explicit the two noise sources unique to this problem: permutation-sampling error and valuation noise from bandit feedback. With bounded Shapley values in M2, MURaS achieves fairness regret M3.
The main algorithm, K-SVFair-FBF, interleaves exploration and exploitation. After round-robin initialization, each round computes Hoeffding-based upper confidence bounds on estimated M4-Shapley values, sets selection probabilities proportional to these optimistic estimates, samples a size-M5 subset via randomized rounding (ensuring unbiased marginal inclusion probabilities), and updates estimates using Monte Carlo permutations combined with M6 repeated pulls for stochastic smoothing. The main theorem states that with probability at least M7, the cumulative fairness regret satisfies
M8
with the smoothing parameter set as M9. This bound sits between the Vt(St)0 lower bound for pure reward maximization under full-bandit feedback and the Vt(St)1 rate of the exploration-separated approach, with the gap attributable to the extra Monte Carlo approximation noise. Two caveats deserve note: the regret is measured over effective rounds Vt(St)2 scaled by internal sampling steps, so the practical per-round cost is substantial (Vt(St)3 evaluations per round); and the analysis assumes bounded rewards and bounded Vt(St)4-Shapley values.
Experimental evaluation
Experiments cover three settings: a synthetic environment with 20 arms, monotone submodular rewards, and Vt(St)5; FL on CIFAR-10 partitioned non-IID (Dirichlet Vt(St)6) across 100 clients with VGG-16 and Vt(St)7; and SIM on a 534-node Facebook subgraph under the Independent Cascade model with Vt(St)8. Baselines include Fair-CMAB (the sole prior FBF fairness work), ETCG, GAP-E, and, for FL, S-FedAvg, ShapFed, ShapleyFL, BSFL, and CS-UCB.
Two metrics are reported. The merit-to-selection ratio (estimated Shapley value divided by empirical selection frequency) remains approximately constant across arms for K-SVFair-FBF and MURaS, confirming proportionality between merit and selection, while baselines exhibit large ratio variation consistent with winner-takes-all behavior. Fairness regret decreases steadily for K-SVFair-FBF across all three datasets and remains below all baselines, which either converge slowly or plateau at higher regret because they do not update selection probabilities from merit estimates. Notably, on the FL task K-SVFair-FBF also attains equal or better global model accuracy than baselines, indicating that merit-proportional participation does not sacrifice learning performance — plausibly because including low-contribution but diverse clients improves generalization under non-IID data. An ablation over Vt(St)9 identifies O(T2/3)0, O(T2/3)1 as the best stability-to-overhead trade-off.
Limitations and open questions
Several limitations are acknowledged or evident. First, the O(T2/3)2 fairness-regret bound carries a multiplicative dependence on O(T2/3)3 and O(T2/3)4, and no matching lower bound exists; the authors explicitly leave open establishing lower bounds that jointly capture stochastic and approximation noise. Second, the framework assumes a stationary environment with fixed arms; dynamic arrival and departure of arms, and fairness trade-offs in non-stationary settings, are stated as unresolved directions. Third, the computational cost of O(T2/3)5 internal evaluations per round is significant, and the reported experiments use horizons far shorter than the asymptotic regime in which the O(T2/3)6 versus O(T2/3)7 separation would be distinguishable. Finally, the fairness guarantee concerns selection proportions relative to estimated merit; whether meritocratic fairness conflicts with other desiderata (e.g., reward maximization regret) is not quantified in this work.
Conclusion
This paper extends meritocratic fairness to budgeted combinatorial bandits under full-bandit feedback, where it was previously unaddressed beyond fixed quotas. Its contributions are an axiomatically characterized O(T2/3)8-Shapley value for coalition-size-restricted games, a UCB-style algorithm that adaptively estimates these values while controlling both valuation noise and Monte Carlo noise, a sublinear O(T2/3)9 fairness-regret guarantee, and empirical validation on FL and influence maximization showing that merit-proportional selection can coexist with strong task performance.