Papers
Topics
Authors
Recent
Search
2000 character limit reached

Meritocratic Fairness in Budgeted Combinatorial Multi-armed Bandits via Shapley Values

Published 1 May 2026 in cs.LG, cs.AI, and cs.MA | (2605.00762v1)

Abstract: We propose a new framework for meritocratic fairness in budgeted combinatorial multi-armed bandits with full-bandit feedback (BCMAB-FBF). Unlike semi-bandit feedback, the contribution of individual arms is not received in full-bandit feedback, making the setting significantly more challenging. To compute arm contributions in BCMAB-FBF, we first extend the Shapley value, a classical solution concept from cooperative game theory, to the KK-Shapley value, which captures the marginal contribution of an agent restricted to a set of size at most KK. We show that KK-Shapley value is a unique solution concept that satisfies Symmetry, Linearity, Null player, and efficiency properties. We next propose K-SVFair-FBF, a fairness-aware bandit algorithm that adaptively estimates KK-Shapley value with unknown valuation function. Unlike standard bandit literature on full bandit feedback, K-SVFair-FBF not only learns the valuation function under full feedback setting but also mitigates the noise arising from Monte Carlo approximations. Theoretically, we prove that K-SVFair-FBF achieves O(T<sup>3/4)O(T<sup>{3/4}) regret bound on fairness regret. Through experiments on federated learning and social influence maximization datasets, we demonstrate that our approach achieves fairness and performs more effectively than existing baselines.

Summary

  • The paper introduces K-Shapley values, an axiomatic merit measure for size-limited coalitions, and defines fairness as selecting arms in proportion to their estimated contributions.
  • The paper’s K-SVFair-FBF algorithm combines UCB exploration, Monte Carlo permutations, repeated sampling, and randomized rounding to achieve high-probability fairness regret of approximately Õ(T^{3/4}KM^{1/2}).
  • Experiments in synthetic bandits, federated learning, and influence maximization show that merit-proportional selection reduces fairness regret and can match or improve task performance versus existing baselines.

Problem setting and motivation

The paper addresses meritocratic fairness in budgeted combinatorial multi-armed bandits under full-bandit feedback (BCMAB-FBF). In this setting, a learner selects a super-arm StS_t of size at most KK from MM base arms and observes only the aggregate stochastic reward Vt(St)V_t(S_t), not per-arm rewards. This contrasts with semi-bandit feedback, where individual arm rewards are visible; under full-bandit feedback, attributing reward to individual arms is fundamentally harder. The authors motivate the problem with federated learning (FL), where a server sees only global model accuracy after aggregating client updates, and social influence maximization (SIM), where only total diffusion spread is observable.

Existing fairness work in CMAB is almost entirely confined to semi-bandit feedback, achieving logarithmic regret via exposure guarantees or linear constraints. Under full-bandit feedback, only one prior work imposes fixed minimum-pull quotas per arm. The authors identify two shortcomings of that approach: quotas are typically unknown a priori, and quota-based fairness ignores merit — an arm's selection frequency should reflect its contribution rather than a static floor. Even ignoring fairness, the best known regret for bandit algorithms under full-bandit feedback is O(T2/3)O(T^{2/3}), so any fairness mechanism must contend with an intrinsically noisy feedback regime.

The K-Shapley value

To define merit without structural assumptions on the reward function, the paper adapts the Shapley value to budgeted coalitions. Since the valuation function VV is defined only on coalitions of size at most KK, classical Shapley axioms cannot be applied directly: marginal contributions to coalitions larger than KK are unobservable, and the grand coalition value V([M])V([M]) is undefined. Existing extensions to restricted cooperative games do not fit — Willson-style values assign zero to infeasible coalitions, Albizuri et al.'s R-games assume grand-coalition feasibility, and graph-restricted games tie feasibility to topology rather than cardinality.

The proposed KK-Shapley value KK0 averages player KK1's marginal contributions over all uniformly random KK2-sized coalitions containing KK3, weighting each by the standard Shapley combinatorial factor within that coalition. The authors prove it satisfies Symmetry, Linearity, Null Player, and a modified K-efficiency axiom requiring

KK4

which reduces to classical efficiency when KK5. A uniqueness theorem establishes that the KK6-Shapley value is the only solution concept satisfying these four axioms, proved by extending carrier-game arguments to KK7-restricted games. Meritocratic fairness is then defined as a policy whose selection probabilities are proportional to KK8-Shapley values, with fairness regret measured as cumulative KK9 deviation from the ideal policy MM0.

Algorithms and theoretical guarantees

The paper first analyzes MURaS (Meritocratic Uniform Random Sampling), an explore-then-commit baseline that estimates MM1-Shapley values via Monte Carlo permutations with repeated stochastic evaluations, then commits to the meritocratic policy. Its analysis makes explicit the two noise sources unique to this problem: permutation-sampling error and valuation noise from bandit feedback. With bounded Shapley values in MM2, MURaS achieves fairness regret MM3.

The main algorithm, K-SVFair-FBF, interleaves exploration and exploitation. After round-robin initialization, each round computes Hoeffding-based upper confidence bounds on estimated MM4-Shapley values, sets selection probabilities proportional to these optimistic estimates, samples a size-MM5 subset via randomized rounding (ensuring unbiased marginal inclusion probabilities), and updates estimates using Monte Carlo permutations combined with MM6 repeated pulls for stochastic smoothing. The main theorem states that with probability at least MM7, the cumulative fairness regret satisfies

MM8

with the smoothing parameter set as MM9. This bound sits between the Vt(St)V_t(S_t)0 lower bound for pure reward maximization under full-bandit feedback and the Vt(St)V_t(S_t)1 rate of the exploration-separated approach, with the gap attributable to the extra Monte Carlo approximation noise. Two caveats deserve note: the regret is measured over effective rounds Vt(St)V_t(S_t)2 scaled by internal sampling steps, so the practical per-round cost is substantial (Vt(St)V_t(S_t)3 evaluations per round); and the analysis assumes bounded rewards and bounded Vt(St)V_t(S_t)4-Shapley values.

Experimental evaluation

Experiments cover three settings: a synthetic environment with 20 arms, monotone submodular rewards, and Vt(St)V_t(S_t)5; FL on CIFAR-10 partitioned non-IID (Dirichlet Vt(St)V_t(S_t)6) across 100 clients with VGG-16 and Vt(St)V_t(S_t)7; and SIM on a 534-node Facebook subgraph under the Independent Cascade model with Vt(St)V_t(S_t)8. Baselines include Fair-CMAB (the sole prior FBF fairness work), ETCG, GAP-E, and, for FL, S-FedAvg, ShapFed, ShapleyFL, BSFL, and CS-UCB.

Two metrics are reported. The merit-to-selection ratio (estimated Shapley value divided by empirical selection frequency) remains approximately constant across arms for K-SVFair-FBF and MURaS, confirming proportionality between merit and selection, while baselines exhibit large ratio variation consistent with winner-takes-all behavior. Fairness regret decreases steadily for K-SVFair-FBF across all three datasets and remains below all baselines, which either converge slowly or plateau at higher regret because they do not update selection probabilities from merit estimates. Notably, on the FL task K-SVFair-FBF also attains equal or better global model accuracy than baselines, indicating that merit-proportional participation does not sacrifice learning performance — plausibly because including low-contribution but diverse clients improves generalization under non-IID data. An ablation over Vt(St)V_t(S_t)9 identifies O(T2/3)O(T^{2/3})0, O(T2/3)O(T^{2/3})1 as the best stability-to-overhead trade-off.

Limitations and open questions

Several limitations are acknowledged or evident. First, the O(T2/3)O(T^{2/3})2 fairness-regret bound carries a multiplicative dependence on O(T2/3)O(T^{2/3})3 and O(T2/3)O(T^{2/3})4, and no matching lower bound exists; the authors explicitly leave open establishing lower bounds that jointly capture stochastic and approximation noise. Second, the framework assumes a stationary environment with fixed arms; dynamic arrival and departure of arms, and fairness trade-offs in non-stationary settings, are stated as unresolved directions. Third, the computational cost of O(T2/3)O(T^{2/3})5 internal evaluations per round is significant, and the reported experiments use horizons far shorter than the asymptotic regime in which the O(T2/3)O(T^{2/3})6 versus O(T2/3)O(T^{2/3})7 separation would be distinguishable. Finally, the fairness guarantee concerns selection proportions relative to estimated merit; whether meritocratic fairness conflicts with other desiderata (e.g., reward maximization regret) is not quantified in this work.

Conclusion

This paper extends meritocratic fairness to budgeted combinatorial bandits under full-bandit feedback, where it was previously unaddressed beyond fixed quotas. Its contributions are an axiomatically characterized O(T2/3)O(T^{2/3})8-Shapley value for coalition-size-restricted games, a UCB-style algorithm that adaptively estimates these values while controlling both valuation noise and Monte Carlo noise, a sublinear O(T2/3)O(T^{2/3})9 fairness-regret guarantee, and empirical validation on FL and influence maximization showing that merit-proportional selection can coexist with strong task performance.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.