Papers
Topics
Authors
Recent
Search
2000 character limit reached

Online Resource Allocation with Continuous Random Consumption: Regret under Degeneracy

Published 2 Jul 2026 in cs.LG | (2607.02196v2)

Abstract: We study online resource allocation when both rewards and consumption sizes may be continuously distributed. Requests arrive sequentially and must be accepted or rejected irrevocably under fixed resource capacities. Each request belongs to one of finitely many observable types; conditional on an observable request type, both the reward and the scalar size are random, and the realized size scales a fixed type-specific resource-consumption vector. The model allows the deterministic fluid relaxation to be degenerate. We show that additive regret is governed by the size-weighted mass of requests whose value-to-size ratios lie near the active acceptance cutoffs. We formalize this quantity through an active weighted-mass exponent p. When p > 1, this cutoff mass is thin, and the problem is genuinely hard: every online policy must incur regret of order at least T<sup>1/2</sup>1/(2p)T<sup>{1/2</sup> - 1/(2p)}, and this holds for every p > 1. A sample-path marginal policy matches this lower bound up to polylogarithmic factors; and when p = 1, so that the mass grows linearly near the cutoff, it attains O((logT)<sup>2)O((\log T)<sup>2) regret. For example, if the size and the value-to-size ratio are independent and uniformly distributed, then p = 1; if instead the size and the reward are independent and uniformly distributed, then p = 2. Thus the policy achieves o(T)o(\sqrt{T}) regret throughout this regularity class without any fluid non-degeneracy assumption, allowing both primal degeneracy and dual non-uniqueness.

Authors (1)

Summary

  • The paper's main contribution reveals that the active weighted-mass exponent governs tight regret bounds for online allocation algorithms.
  • It shows that sufficient mass near cutoffs yields polylogarithmic regret, while thin mass under degeneracy causes a polynomial regret lower bound.
  • The proposed sample-path marginal (SPM) policy avoids reliance on unstable dual prices, ensuring robust performance even in degenerate settings.

Online Resource Allocation with Continuous Random Consumption: Regret under Degeneracy

Problem Setting and Motivation

The paper investigates the online resource allocation problem under significant generality: both rewards and consumption sizes are continuously distributed random variables, requests arrive sequentially and reveal their characteristics upon arrival, and each must be irrevocably accepted or rejected, observing only type, reward, and realized size. Each request consumes a (random) scalar of a type-specific deterministic resource vector, and all resources have fixed capacities over a finite horizon.

The core objective is to analyze the additive regret of online algorithms relative to the offline fractional hindsight optimum, which selects the optimal subset of requests after observing the full arrival sequence. Of particular interest is how the structure of joint reward and resource consumption distributions — and especially degeneracies in the associated fluid LP relaxation — impact the achievable regret rates. The focus is on identifying the mechanisms driving worst-case regret and on constructing online policies that match the inherent complexity induced by degeneracy and thin mass near cutoffs.

Model and Prior Results

The model generalizes classical online allocation and stochastic knapsack settings, allowing for a finite set of observable types, i.i.d. arrivals, and continuous joint distributions for rewards and sizes. For each type, resource consumption is given by a fixed vector scaled by a random size; rewards and sizes can exhibit arbitrary dependence.

A standard approach to the online allocation problem is to use dual prices derived from periodically re-solving the fluid LP relaxation, or variants thereof, to guide accept/reject decisions. For finite-support or non-degenerate (interior) fluid relaxations, sublinear (often polylogarithmic or even constant) additive regret bounds have been established for these methods [e.g., (Mandal et al., 2018, Fan et al., 2019)]. However, these guarantees crucially rely on strong regularity—uniqueness or stability of the dual solution, absence of critical degeneracy, or adequate mass near accept/reject cutoffs.

When degeneracy is present, prior work often observes only at most a mild increase in regret—typically to an extra logarithmic factor—provided that consumption is deterministic (conditioned on type) or reward distributions are sufficiently regular [see JiangMaZhang2025Degeneracy, BesbesEtAl2024]. However, the case when both rewards and consumption are continuously and jointly random, and the fluid relaxation is degenerate, was not previously resolved.

Main Contributions

The paper identifies a sharp dichotomy in regret behavior, parameterized by the growth of the size-weighted value-to-size ratio measure near active cutoffs, formalized by an active weighted-mass exponent p1\mathfrak{p} \geq 1. The key results are:

  • If p=1\mathfrak{p} = 1 (i.e., sufficient mass near cutoff): Additive regret of the online policy is O((logT)2)O((\log T)^2). This regime includes, for example, the case where the ratio R=V/βR = V/\beta is independent of β\beta and uniformly distributed.
  • If p>1\mathfrak{p} > 1 (thin mass near cutoff): Every online policy incurs regret at least Ω(T1/21/(2p))\Omega\left(T^{1/2-1/(2\mathfrak{p})}\right), and the constructed sample-path marginal (SPM) policy matches this lower bound (up to polylog terms).

Critically, thinness of mass arises naturally under degeneracy: when the critical cutoff for RR sits at a corner of the support, which may only be approached as rewards and sizes simultaneously hit endpoints, the size-weighted measure may grow only as xpx^{\mathfrak{p}}, p>1\mathfrak{p} > 1. For example, if p=1\mathfrak{p} = 10 and p=1\mathfrak{p} = 11 are independent and uniform, then p=1\mathfrak{p} = 12, yielding regret of p=1\mathfrak{p} = 13.

These results do not require any non-degeneracy or dual uniqueness assumptions of the fluid relaxation—meaning both primal degeneracy and dual non-uniqueness are explicitly allowed.

The Sample-Path Marginal Policy

The proposed SPM policy eschews using a fixed dual price; instead, it accepts a request if its reward exceeds the sample-path expected marginal value of the capacity it would consume, as computed from the future paths of i.i.d. arrivals, thereby averaging over all possible dual prices along the entire relevant capacity trajectory. This approach remains well-defined and stable under degeneracy—where selected dual prices can jump discontinuously or be non-unique.

Analytical Framework

The core technical insight is that the regret per period is governed not just by the variability of cutoffs but by the product of cutoff width and the size-weighted mass swept by this width near the cutoff. Even if empirical cutoffs can be far apart in degenerate cases, the loss can be small provided the resource mass scanned is minimal (as encoded by the exponent p=1\mathfrak{p} = 14). The analysis leverages a detailed convex analysis, measure decomposition at endpoints (introducing the notion of "branch contact"), and LP stability (Hoffman's lemma) arguments to formalize this mechanism.

The per-stage analysis is reduced to bounding Jensen gaps for convex functions, which ultimately depend on controlling the product of (i) the width between empirical cutoffs arising from two possible future paths and (ii) the size-weighted measure of resources having value-to-size ratio within this width. Intricate endpoint analysis is conducted when the cutoff coincides with a corner of the support—requiring Hardy-type weighted integral bounds over the "branch" of sizes that can approach the cutoff.

A matching lower bound is provided by exhibiting instances (critical-capacity, endpoint-pinned degeneracy with vanishing resource mass near cutoff) for which no policy can do better than p=1\mathfrak{p} = 15 additive regret.

Summary of Results

Strong numerical/implied results include:

Value of Exponent p=1\mathfrak{p} = 16 Regret Rate of SPM Policy Regime Characteristics
p=1\mathfrak{p} = 17 p=1\mathfrak{p} = 18 Regular, sufficient size-weighted mass at (interior) cutoff
p=1\mathfrak{p} = 19 O((logT)2)O((\log T)^2)0 Endpoint-critical, thin resource mass (e.g., reward and size independent uniform: O((logT)2)O((\log T)^2)1)

These rates are tight: policies cannot be improved beyond these bounds without additional assumptions on the distributions or the non-degeneracy of the fluid relaxation.

Notably, the phenomenon is critical-capacity specific: at fixed non-critical capacities (i.e., interior cutoffs), the local exponent may remain at one and the regret rate returns to polylogarithmic, even for the same arrival distributions.

Implications and Open Directions

The work rigorously characterizes the previously unexplored cost of degeneracy in online resource allocation with continuous random consumption, revealing that polylogarithmic regret is lost only when the joint distribution conspires to make resource mass near cutoff thin—typically at corners induced by critical degeneracy. In practical terms, this highlights the limits of fluid-based or certainty-equivalent threshold policies under random, continuous consumption, and guides the design and analysis of robust policies in such degenerate regimes.

This analysis circumvents the use of dual prices, instead exploiting averaged marginal values, thus producing policies resilient to instability in dual solutions; this offers direct implications for designing allocation algorithms that avoid "chasing moving thresholds" under degeneracy.

Future research directions include:

  • Developing efficient variants of the SPM policy that avoid full simulation of sample-path marginals, potentially leveraging approximate dynamic or empirical-marginal approaches.
  • Extending the analysis to fully random (not just scalar-multiplied) consumption vectors, which would require more refined measure-theoretic and geometric LP sensitivity techniques.
  • Exploring unknown distribution regimes (addressed in a companion note) and evaluating practical estimation and learning schemes under minimal regularity.

Conclusion

This paper provides a definitive account of regret under degeneracy in online allocation with continuous random reward and consumption, precisely capturing the polynomial price incurred when joint randomness and degeneracy intersect to create thin mass near acceptance cutoffs, while showing polylogarithmic regret remains achievable when sufficient size-weighted mass persists. These findings unify and extend several lines of prior work, and delineate the precise regime where degeneracy is costly in online stochastic optimization (2607.02196).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.