Papers
Topics
Authors
Recent
Search
2000 character limit reached

Improved Algorithms for Nash Welfare in Linear Bandits

Published 30 Jan 2026 in cs.LG | (2601.22969v1)

Abstract: Nash regret has recently emerged as a principled fairness-aware performance metric for stochastic multi-armed bandits, motivated by the Nash Social Welfare objective. Although this notion has been extended to linear bandits, existing results suffer from suboptimality in ambient dimension dd, stemming from proof techniques that rely on restrictive concentration inequalities. In this work, we resolve this open problem by introducing new analytical tools that yield an order-optimal Nash regret bound in linear bandits. Beyond Nash regret, we initiate the study of pp-means regret in linear bandits, a unifying framework that interpolates between fairness and utility objectives and strictly generalizes Nash regret. We propose a generic algorithmic framework, FairLinBandit, that works as a meta-algorithm on top of any linear bandit strategy. We instantiate this framework using two bandit algorithms: Phased Elimination and Upper Confidence Bound, and prove that both achieve sublinear pp-means regret for the entire range of pp. Extensive experiments on linear bandit instances generated from real-world datasets demonstrate that our methods consistently outperform the existing state-of-the-art baseline.

Summary

  • The paper presents a two-phase algorithm, FairLinBandit, that achieves Nash regret $O(\sigma d\log T/\sqrt T)$ in linear bandits, improving prior $\widetilde O(d^{5/4}/\sqrt T)$ Nash regret by removing the reliance on estimate-dependent
  • The algorithm employs a data-adaptive stopping rule, ensuring fairer Nash welfare with an efficient, effective exploration-exploitation strategy. Tested methods included phase exploration (Phase I), optimizing the learner's estimation of values, and a modified LinUCB or LinPE (Phase II).
  • Variable regimes for $p$-means regret updated. $p=0$: Standard Nash regret $B_p$ is equivalent to $p\in[-0.5,0]$ for sub-ball extrapolation. Beyond Nash: Regime $p \leq -1$ degradation asymptotically with $|p|$ for fair-Pareto optimization.

Background and problem

Nash regret evaluates a bandit policy by the gap between the optimal expected reward and the geometric mean of expected per-round rewards, replacing the arithmetic mean of standard regret. Because the geometric mean penalizes rounds with very low reward, it encodes a fairness criterion grounded in Nash Social Welfare. Prior work established (near-)optimal Nash and pp-means regret bounds for finite-armed stochastic bandits [barman2023fairness; krishna2025p; sarkar2025revisiting], but the linear bandit extension by Sawarni et al. achieved only O~(d5/4/T)\widetilde O(d^{5/4}/\sqrt T) Nash regret against an Ω(d/T)\Omega(d/\sqrt T) lower bound inherited from average-regret minimization in infinite-armed linear bandits [dani2008stochastic]. That suboptimality was traced to estimate-dependent "Nash confidence bounds" (NCBs), which require restrictive multiplicative concentration inequalities and non-negative sub-Poisson rewards. The paper resolves this open problem and initiates pp-means regret minimization for linear bandits.

The FairLinBandit meta-algorithm

FairLinBandit is a two-phase meta-algorithm wrapped around any optimistic linear bandit strategy. Phase I interleaves two geometric sampling schemes with equal probability: arms drawn from a D-optimal design λ∗\lambda^* (which by Kiefer–Wolfowitz also minimizes maximum predictive variance, giving x⊤V−1x≤3d/tx^\top V^{-1}x \le 3d/t), and arms sampled from a distribution induced by the John ellipsoid center of conv(X)\mathrm{conv}(\mathcal X), guaranteeing expected reward at least ⟨x∗,θ∗⟩/(d+1)\langle x^*,\theta^*\rangle/(d+1) so the geometric mean does not collapse.

The central methodological contribution is a data-adaptive stopping rule for Phase I: exploration continues until the least-squares estimate θ^\widehat\theta satisfies a condition of the form

48σ2d2log⁡T(max⁡x⟨x,θ^⟩)2<t≤900p2σ2d2log⁡T(max⁡x⟨x,θ^⟩−48σ2d2log⁡T/t)2,\frac{48\sigma^2 d^2\log T}{(\max_x \langle x,\widehat\theta\rangle)^2} < t \le \frac{900 p^2\sigma^2 d^2\log T}{(\max_x \langle x,\widehat\theta\rangle - \sqrt{48\sigma^2 d^2\log T/t})^2},

implemented via doubling epochs with persistent sufficient statistics. This forces Phase I to run for O~(d5/4/T)\widetilde O(d^{5/4}/\sqrt T)0 rounds, ensuring that standard UCB confidence widths O~(d5/4/T)\widetilde O(d^{5/4}/\sqrt T)1 never exceed the optimal reward — precisely the failure mode that previously forced reliance on NCBs. Phase II then runs LinUCB or Linear Phased Elimination (LinPE) initialized with the Phase I statistics. LinPE restricts to a surviving near-optimal arm set and re-solves D-optimal designs over doubling episodes; it discards history between episodes but has lower per-round cost than LinUCB.

Regret guarantees

The main theorem shows that under norm-boundedness, sub-Gaussian rewards, and non-negative means, FairLinBandit with either instantiation achieves Nash regret O~(d5/4/T)\widetilde O(d^{5/4}/\sqrt T)2, matching the O~(d5/4/T)\widetilde O(d^{5/4}/\sqrt T)3 lower bound up to logarithmic factors. This improves the previous best O~(d5/4/T)\widetilde O(d^{5/4}/\sqrt T)4 and resolves the open question posed by Sawarni et al. The proof splits the Nash Social Welfare product across phases: Phase I rewards are lower bounded via the John ellipsoid guarantee, while Phase II exploits arm-independent UCB widths together with Cauchy–Schwarz control of the cumulative elliptical norms.

The framework extends to O~(d5/4/T)\widetilde O(d^{5/4}/\sqrt T)5-means regret for all O~(d5/4/T)\widetilde O(d^{5/4}/\sqrt T)6:

Regime O~(d5/4/T)\widetilde O(d^{5/4}/\sqrt T)7-means regret bound
O~(d5/4/T)\widetilde O(d^{5/4}/\sqrt T)8 O~(d5/4/T)\widetilde O(d^{5/4}/\sqrt T)9
Ω(d/T)\Omega(d/\sqrt T)0 Ω(d/T)\Omega(d/\sqrt T)1
Ω(d/T)\Omega(d/\sqrt T)2 Ω(d/T)\Omega(d/\sqrt T)3

For Ω(d/T)\Omega(d/\sqrt T)4 the bound is order-optimal since power-mean regret dominates average regret. For negative Ω(d/T)\Omega(d/\sqrt T)5, the dependence on Ω(d/T)\Omega(d/\sqrt T)6 degrades exponentially in Ω(d/T)\Omega(d/\sqrt T)7 unless Ω(d/T)\Omega(d/\sqrt T)8, an explicit fairness–performance trade-off the authors characterize as a "no free lunch" phenomenon. Notably, the analysis covers general Ω(d/T)\Omega(d/\sqrt T)9-sub-Gaussian rewards rather than only non-negative sub-Poisson rewards; the appendix argues that the prior NCB-based analysis breaks down even for non-negative sub-Gaussian rewards, because substituting the implied sub-Poisson parameter makes their confidence widths estimate-independent and invalidates the argument.

Compared with the pp0-armed results of Sarkar et al., the bounds carry an extra factor of roughly pp1 when specialized (pp2), attributed to the harder infinite-arm setting where rewards are coupled through a shared parameter.

Reduction framework

The stopping-rule construction yields a plug-and-play reduction: given (i) an arm-independent confidence width bound after Phase I, (ii) a time-uniform per-round regret bound for Phase II, and (iii) a suitably modified termination condition, any average-regret algorithm can be lifted to Nash regret minimization. Plugging in SupLinUCB for finitely many arms yields Nash regret pp3, matching known lower bounds. Plugging in LinTS inherits its known suboptimal average regret, producing pp4 Nash regret. The authors conjecture the reduction extends to generalized linear (e.g., logistic) rewards, but this remains unverified.

Experiments

Experiments use linear bandit instances derived from MSLR-WEB10K (908 arms, reduced to pp5 via PCA) and Yahoo! Learning to Rank Challenge data, with Gaussian noise pp6. Both FairLinPE and FairLinUCB converge faster and more stably than LinNash on both datasets; at pp7, LinNash exhibits instability attributable to its wide, dimension-sensitive confidence intervals. Across pp8, FairLinUCB attains lower pp9-means regret than FairLinPE, consistent with its tighter per-round estimates. Ablations confirm that regret increases as λ∗\lambda^*0 decreases and grows roughly linearly with λ∗\lambda^*1, corroborating the theoretical dependence. Runtime comparisons show a clear trade-off: FairLinPE is substantially faster than FairLinUCB (e.g., about 3,944s versus 58,201s at λ∗\lambda^*2 on MSLR-WEB10K) while still outperforming LinNash in both speed and regret.

Limitations and open questions

Several caveats are explicit in the paper. First, the tightness of the λ∗\lambda^*3-means regret upper bound for λ∗\lambda^*4 is unresolved; whether the exponential-in-λ∗\lambda^*5 dependence is necessary is stated as an open question. Second, the bounds depend on the optimal reward through the Phase I length λ∗\lambda^*6, so instances with small λ∗\lambda^*7 require long exploration before exploitation begins. Third, the reduction framework's extension beyond linear rewards (e.g., logistic bandits) is conjectural. Fourth, the LinPE analysis assumes episode lengths exceed λ∗\lambda^*8, implicitly requiring λ∗\lambda^*9; the authors note this is benign since even minimax rates are vacuous below x⊤V−1x≤3d/tx^\top V^{-1}x \le 3d/t0, but it is nonetheless an assumption. Finally, the empirical evaluation uses PCA-reduced, Lasso-fitted surrogates of real datasets rather than online deployments, so the practical fidelity of the constructed instances is approximate.

Conclusion

This work resolves the open problem of order-optimal Nash regret in linear bandits by replacing estimate-dependent confidence bounds with UCB-style intervals enabled by a data-adaptive exploration stopping rule, achieving x⊤V−1x≤3d/tx^\top V^{-1}x \le 3d/t1 Nash regret and initiating the study of x⊤V−1x≤3d/tx^\top V^{-1}x \le 3d/t2-means regret in this setting. The resulting meta-algorithmic framework is modular, supports sub-Gaussian rewards, and is validated empirically against the prior state of the art. The principal open issues are matching lower bounds for strongly negative x⊤V−1x≤3d/tx^\top V^{-1}x \le 3d/t3 and extending the reduction beyond linear reward models.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.