---
title: Learning Curves for Revenue Maximization
url: https://www.emergentmind.com/topics/learning-curves-for-revenue-maximization
type: topic
---

# Learning Curves for Revenue Maximization

Searching arXiv for the core paper and closely related work on learning curves for revenue maximization.
arXiv search: "On the Learning Curves of Revenue Maximization" 2604.26922
Learning curves for revenue maximization quantify how rapidly a pricing or mechanism-design procedure approaches its benchmark revenue as information accumulates. In the single-item, single-buyer posted-pricing model, the canonical object is the expected revenue gap
$$
\epsilon_n(t_n,D):=\operatorname{opt}_D-\mathbb{E}[\operatorname{rev}_D(t_n)],
$$
viewed as a function of the sample size $n$ for a fixed valuation distribution $D$ [2604.26922]. In adjacent literatures, the same idea is instantiated through relative regret against a fluid benchmark, dynamic regret against a time-varying oracle, Bayesian regret across episodes, pricing-query complexity, and sample complexity for approximate optimality. The resulting theory shows that the shape of the curve depends on the benchmark, the feedback model, the regularity of demand, whether the optimal price is attained, and whether the seller faces inventory, budget, or strategic-response constraints.

## 1. Formalizations and benchmark choices

The fixed-distribution formulation isolates a single distribution $D$ and studies the sequence $\epsilon_n(t_n,D)$ for an algorithm $(t_n)$. In this framework, Bayes-consistency means
$$
\lim_{n\to\infty}\mathbb{E}[\operatorname{rev}_D(t_n)]=\operatorname{opt}_D
$$
for every $D$ in the class under study, with the convention that if $\operatorname{opt}_D=\infty$, then $\mathbb{E}[\operatorname{rev}_D(t_n)]\to\infty$ [2604.26922]. The same paper distinguishes a PAC upper bound, which controls the worst-case envelope uniformly over a class of distributions, from a universal learning rate, which allows the constants to depend on the fixed underlying distribution. This distinction is central because fixed-distribution learning curves can be much sharper than worst-case PAC rates.

Online and dynamic models replace the sample-size axis by a time horizon and define learning curves through regret. In personalized dynamic pricing with an inventory constraint, regret is
$$
R_n^{\pi}(T,c)\equiv 1-\frac{J_n^{\pi}(T,c)}{J_n^D(T,c)},
$$
where $J_n^D(T,c)$ is the fluid benchmark under proportional scaling of demand and capacity [1812.09234]. Under nonstationarity with one-point feedback, the benchmark is dynamic:
$$
\operatorname{D\text{-}Reg}(T)=\sum_{t=1}^T r_t(x_t^*)-\mathbb{E}\!\left[\sum_{t=1}^T r_t(\tilde x_t)\right],
$$
with $x_t^*\in\arg\max_{x\in K} r_t(x)$ and path variation
$$
V_T=\max_{x_t^*\in\Theta_t^*}\sum_{t=2}^T \|x_t^*-x_{t-1}^*\|
$$
as the nonstationarity budget [2605.21263]. In pricing-query models, the learning curve is expressed by the achievable revenue gap after $T$ pricing queries, typically denoted $\operatorname{gap}(T)$, while in episodic revenue management with unknown time-varying demand, the benchmark is Bayesian regret relative to the clairvoyant dynamic-programming policy over $S$ episodes of length $T$ [2111.03158][2405.04910].

These benchmark choices are not interchangeable. A fixed-distribution curve measures convergence to optimal revenue for one environment; a regret curve measures cumulative shortfall under sequential decision-making; a pricing-query curve measures how much can be inferred from binary accept/reject observations. Much of the contemporary literature can be read as a comparison of these benchmark families.

## 2. Distribution-dependent sample-size regimes

The sharpest current characterization of fixed-distribution learning curves is for the basic single-item, single-buyer posted-pricing model [2604.26922]. The central structural distinction is whether the optimal revenue is attained at a finite price.

| Regime | Learning-curve rate | Source |
|---|---|---|
| $\operatorname{opt}_D<\infty$, not attained at finite price | arbitrarily slow | [2604.26922] |
| $\operatorname{opt}_D$ attained at finite $p^*$ | essentially $1/\sqrt{n}$ up to logarithmic factors | [2604.26922] |
| bounded support | optimal $o(n^{-1/2})$ universal rate | [2604.26922] |
| closed discrete support | $e^{-o(n)}$ | [2604.26922] |
| finite support | $e^{-n}$ | [2604.26922] |

For unrestricted distributions, there exists a Bayes-consistent algorithm for all valuation distributions on $\mathbb{R}_+$: capped ERM with cap $\log n$ achieves $\mathbb{E}[\operatorname{rev}_D(t_n)]\to \operatorname{opt}_D$ even when $\operatorname{opt}_D=\infty$ [2604.26922]. However, if $\operatorname{opt}_D<\infty$ but no finite price attains it, then convergence can be arbitrarily slow. The lower bound holds in a strong sense: for any algorithm $(t_n)$ and any rate function $\phi(n)\downarrow 0$, there exists a fixed distribution $D$ such that infinitely often
$$
\epsilon_n(t_n,D)\ge c\cdot \phi(n)
$$
for some constant $c$ [2604.26922]. The heavy-tail example
$$
F(v)=1-\frac{1}{v+1},\qquad R(p)=\frac{p}{p+1}\to 1
$$
illustrates the regime $\operatorname{opt}_D<\infty$ without attainment [2604.26922].

When the optimal revenue is achieved at a finite price $p^*$, the universal rate becomes essentially $1/\sqrt{n}$. More precisely, for any rate $R(n)\in\omega(n^{-1/2})$ there is an algorithm that learns the class at universal rate $R(n)$, whereas for any $R(n)\in o(n^{-1/2})$ no algorithm can learn the class universally at that rate [2604.26922]. On bounded supports, ERM improves this to an optimal $o(n^{-1/2})$ universal rate by a localized Bernstein analysis. On closed discrete supports, structured ERM attains $e^{-o(n)}$ rates, while vanilla ERM is not Bayes-consistent: there exists a closed discrete support on which ERM incurs a constant expected revenue gap along an infinite subsequence of sample sizes [2604.26922]. On finite supports, ERM achieves
$$
\epsilon_n(t_n,D)\le c_1(D)e^{-c_2(D)n},
$$
and no faster-than-exponential universal rate is possible on nontrivial finite supports [2604.26922].

A recurring conclusion is that PAC-style worst-case envelopes obscure these shapes. The fixed-distribution view separates heavy-tail nonattainment, finite-price attainability, bounded support, closed discrete support, and finite support into genuinely different rate classes.

## 3. Inventory, resource constraints, and contextual decision-making

A large segment of the literature studies learning curves in revenue maximization under operational constraints rather than pure posted-pricing estimation. In personalized dynamic pricing with one inventory resource and $M$ observable customer types, the seller learns a single dual shadow price $z^*$ instead of learning all type-specific demand functions in full [1812.09234]. The fluid benchmark is
$$
J^D(T,c)\equiv \max_{\{p_m(t)\}} \sum_{m=1}^M\int_0^T p_m(t)d_m(p_m(t))\,dt
\quad \text{s.t.}\quad
\sum_{m=1}^M\int_0^T d_m(p_m(t))\,dt\le c,
$$
and the dual variable $z^*$ satisfies
$$
z^*=\arg\min_{z\ge 0}\Big\{cz+T\sum_{m=1}^M \mathcal{R}_m(z)\Big\},
\qquad
p_m^*=\mathcal{P}_m(z^*).
$$
The primal-dual learning algorithm achieves the dimension-free regret rate
$$
R_n^{\pi}(T,c)\le C(\log n)^{2+\delta}n^{-1/2},
$$
with the exponent independent of the number of types $M$ [1812.09234]. Under sufficient capacity, the final phase uses one price per type; under insufficient capacity, it uses two prices per type and an interpolation parameter $\theta$ to pin aggregate sales near $nc$.

In the older single-product, limited-inventory model with unknown regular demand, the dynamic pricing algorithm of Besbes and Zeevi uses shrinking price intervals and function-value estimation rather than parametric identification [1101.4681]. In the size-$n$ scaling regime, it achieves
$$
\sup_{\lambda\in\Gamma} R_n^\pi(x,T;\lambda)\le \frac{C(\log n)^{4.5}}{\sqrt{n}},
$$
while a lower bound shows that no admissible policy can beat order $1/\sqrt{n}$ up to logarithmic factors [1101.4681]. The paper interprets this as closing the gaps between parametric and non-parametric learning and between a post-price mechanism and a customer-bidding mechanism.

Constraint-coupled contextual revenue maximization introduces a different learning curve. In dual mirror descent with unknown model parameter $\theta^*$, the agent observes i.i.d. contexts $w^t$, chooses $z^t$, receives revenue $f(z^t;\theta^*,w^t)$, and incurs costs $c(z^t;\theta^*,w^t)$ subject to both lower and upper average-cost bounds [2104.09750]. With known $\theta^*$, regret and lower-bound violations are both $O(\sqrt{T})$. With unknown $\theta^*$, the decomposition
$$
\operatorname{Regret}(A|P)\le \Delta_{\mathrm{DM}}+\Delta_{\mathrm{Learn}}
$$
adds an estimation term that depends on
$$
\mathbb{E}\!\left[\sum_{t=1}^{\tau_A}\|\theta^*-\theta^t\|_\theta\right].
$$
If $\|\theta^t-\theta^*\|_\theta=O(t^{-1/2})$, the overall regret remains $O(\sqrt{T})$; if it decays as $O(1/t)$, the bound becomes $O(\sqrt{T}+\log T)$ [2104.09750].

These results share a common pattern: the learning curve is not only a function of statistical difficulty, but also of how a low-dimensional structure—such as a dual variable, a shrinking interval, or a dual feasibility certificate—compresses the constrained control problem.

## 4. Nonstationarity and time-varying demand

When demand changes over time, learning curves acquire an explicit dependence on a variation budget. Under one-point feedback in a convex price domain $K\subset\mathbb{R}_+^d$, mirror ascent with the estimator
$$
G_t=\frac{\phi_t(\tilde x_t)}{\delta}\,\hat U_t,\qquad \tilde x_t=x_t+\delta \tilde U_t,
$$
and periodic restarting yields a static regret bound
$$
\max_{x\in\Theta}\sum_{t=1}^\tau r_t(x)-\mathbb{E}\!\left[\sum_{t=1}^\tau r_t(\tilde x_t)\right]
\le \frac{B}{\eta}+\frac{\tau\eta}{\alpha}\Big(C_1^2\delta^{2p}+\frac{C_2}{2}\delta^{-q}+C_r\Big)+\tau C_1\delta^p\sqrt d+\frac{\tau L C_u}{2}\delta^2
$$
[2605.21263]. With tuned $\delta$ and $\eta$, this becomes $O(\tau^{(\hat p+q)/(2\hat p+q)})$, and for spherical smoothing $(p=q=2)$ it reduces to $O(\tau^{2/3})$. Restarting converts this into a dynamic regret bound
$$
O\!\big(\operatorname{poly}(d)\,T^{(2\hat p+q)/(3\hat p+q)}V_T^{\hat p/(3\hat p+q)}\big),
$$
with the special case
$$
\operatorname{D\text{-}Reg}(T)=O\!\big(\operatorname{poly}(d)\,T^{3/4}V_T^{1/4}\big)
$$
when $p=q=2$ and $V_T$ is known. If $V_T$ is unknown, the bandit-over-bandit meta-layer yields
$$
O\!\big(\operatorname{poly}(d)\,T^{7/9}V_T^{1/3}\big)
$$
for $p=q=2$ [2605.21263].

In the single-buyer binary-feedback model with drifting valuations, the benchmark is first-best revenue
$$
\operatorname{Rev}^*(1{:}T)=\sum_{t=1}^T v_t,
$$
and the regret is
$$
\mathcal{R}_T=\operatorname{Rev}^*(1{:}T)-\sum_{t=1}^T p_t\sigma_t.
$$
Here the nonstationarity parameter is either a fixed $\epsilon$, the average changing rate
$$
\bar\epsilon=\frac{1}{T-1}\sum_{t=1}^{T-1}\epsilon_t,
$$
or
$$
\tilde\epsilon=\sqrt{\frac{1}{T}\sum_{t=1}^T\epsilon_t^2}
$$
in the stochastic known-$\epsilon_t$ case [2106.04689]. The optimal exponents differ between adversarial and stochastic drift:
$$
\frac{\mathcal{R}_T}{T}=\tilde O(\epsilon^{1/2}) \quad \text{adversarial},\qquad
\frac{\mathcal{R}_T}{T}=\tilde O(\epsilon^{2/3}) \quad \text{stochastic},
$$
with matching lower bounds, and the same exponents extend to unknown or dynamic non-increasing $\epsilon_t$ through $\bar\epsilon$ [2106.04689]. The algorithms alternate binary-search localization with exploitation phases and sparse checking rounds.

Episodic revenue management with unknown time-varying demand introduces a further interaction between learning and inventory. With $S$ seasons, $T$ periods per season, a finite price set $\mathcal P$, and a Bayesian prior over time-varying demand parameters, posterior sampling plus LP re-optimization yields
$$
\operatorname{BRegret}(T,S,f,\pi)\le p_M \bar d\Big(S\sqrt{T}+54T\sqrt{SK\log K}\Big),
$$
while a lower bound shows
$$
\operatorname{BRegret}(T,S,f_0,\pi)\ge \Omega\!\big(T\sqrt{SK}\big)
$$
for some prior $f_0$ [2405.04910]. The paper interprets the $T\sqrt{SK\log K}$ term as the unavoidable learning component and the $S\sqrt T$ term as the computational-efficiency cost of using LP rather than exact dynamic programming. Empirically, correlated Gaussian-process priors steepen early learning curves by sharing information across time and price [2405.04910].

Taken together, these models replace a single asymptotic rate by a two-parameter geometry: horizon length controls the accumulation of exploration error, while a variation budget controls how rapidly past information becomes stale.

## 5. Feedback limitations, strategic responses, and information complexity

Learning curves also depend on what information the seller receives. In pricing-query models, the learner posts a price $p_t$, observes only the binary signal
$$
s_t=\operatorname{sign}(p_t-v_t)\in\{-1,+1\},
$$
and seeks a reserve price $r$ that nearly maximizes
$$
R(p)=p(1-F(p))=pQ(p)
$$
[2111.03158]. The resulting query complexity exhibits three regimes:
$$
\tilde\Theta(\epsilon^{-3}) \text{ for general distributions},\qquad
\tilde\Theta(\epsilon^{-2}) \text{ for regular distributions},\qquad
\tilde\Theta(\epsilon^{-2}) \text{ for MHR distributions}.
$$
Equivalently, the revenue-gap learning curve satisfies
$$
\operatorname{gap}(T)\approx \tilde\Theta(T^{-1/3})
$$
for general distributions and
$$
\operatorname{gap}(T)\approx \tilde\Theta(T^{-1/2})
$$
for regular and MHR distributions [2111.03158]. The regular-distribution algorithm relies on the relative flatness property, which rules out hidden interior spikes after probing a constant number of evenly spaced prices. The same paper shows that for regular distributions, learning the reserve is strictly easier than learning the entire distribution in Lévy distance:
$$
\operatorname{PriceCplx}_{\mathrm{Regular},F}(\epsilon)\ge \Omega(\epsilon^{-2.5}),
$$
whereas reserve-price learning requires only $\tilde O(\epsilon^{-2})$ queries [2111.03158].

Patient buyers create a different information geometry. When each buyer has a value-patience type $(v,w)$ and can delay purchase over up to $w$ steps, the revenue class for pure non-increasing price sequences has fat-shattering dimension linear in $w$ [2202.06143]. The offline pure-strategy sample complexity has two regimes: for sample size $m\gtrsim w^3$ the excess error behaves like $O(\sqrt{w/m})$, whereas for $m\lesssim w^3$ it behaves like $O((\log w/m)^{1/3})$ up to logarithmic factors. In online learning, regret against the optimal pure strategy is
$$
O(\sqrt{wT})
$$
after the crossover $T\gtrsim w^3/\log^2 w$, while regret against the optimal mixed strategy is $O(\sqrt{T|V|w})$ for finite support and $O(T^{2/3}w^{4/3})$ in general [2202.06143].

Strategic data generation further modifies the curve. In ERM with endogenous sampling, the “samples” are buyers’ bids, and a coalition of size $m$ can manipulate the learned reserve. The paper formalizes this through the incentive-awareness measure
$$
\Delta_{N,m}^{\mathrm{worst}}
=
\mathbb{E}_{v_{-I}\sim F}\big[\delta_m^{\mathrm{worst}}(v_{-I})\big],
$$
which upper-bounds the expected relative price drop caused by altering $m$ out of $N$ samples [2010.05519]. For guarded ERM,
$$
\Delta_{N,m}^{\mathrm{worst}}=O\!\left(\frac{m\log^3 N}{\sqrt N}\right)
$$
for MHR distributions and
$$
\Delta_{N,m}^{\mathrm{worst}}=O\!\left(D^{8/3}m^{2/3}\frac{\log^2 N}{N^{1/3}}\right)
$$
for distributions supported on $[1,D]$ [2010.05519]. The endogenous-learning curve therefore becomes the sum of the classical statistical error and a manipulation-robustness term:
$$
\epsilon^{\mathrm{sample}}(N)+\Theta(\Delta_{N,m}^{\mathrm{worst}}).
$$

Two-sided learning against a budget- and ROI-constrained buyer supplies a further example of structure-driven learnability. The seller’s fixed-price revenue
$$
\pi(d)=d\sum_{n=1}^N g_n x_{d,n}
$$
is bell-shaped over the price grid: strictly increasing on the non-binding region, flat at $\rho$ on the budget-binding segment, and strictly decreasing on the ROI-binding segment [2107.07725]. An episodic binary search then achieves
$$
\operatorname{Reg}_T
\le
2(\lfloor\log_2 M\rfloor+1)T^{1-\xi+\epsilon}
+\phi(T)
+\frac{(\lfloor\log_2 M\rfloor+1)^2}{2},
$$
where $\xi$ is the buyer’s within-episode adaptivity exponent and $\phi(x)=O(x^{1-\xi})$ [2107.07725]. If the buyer best responds exactly, or uses empirical-distribution advice, then $\xi=1/2$ and the seller’s regret is $\Theta(T^{1/2+\epsilon})$.

## 6. Algorithmic motifs, applications, and open directions

Several recurring design principles organize the field. Capped ERM and structured ERM control tail risk or discrete-support pathologies in fixed-distribution learning [2604.26922]. Primal-dual methods reduce a high-dimensional pricing problem to learning a shadow price of capacity or cost [1812.09234][2104.09750]. Restarting and meta-learning discount stale information under nonstationarity [2605.21263]. Relative flatness enables zoom-in without reconstructing the full value distribution [2111.03158]. Posterior sampling combined with LP re-optimization transforms episodic time-varying revenue management into tractable mean-demand planning [2405.04910].

These motifs extend beyond single posted prices. In airline revenue management with unconstrained capacity and simultaneous pricing of $H$ active flights, the unified objective
$$
U(\pi)=R(\pi)-\eta\cdot \frac{1}{\hat\phi_t\sqrt{\mathcal I_t(\pi)}}
$$
balances current expected revenue against Fisher-information-driven learning quality [2203.11065]. In the reported experiments, with $H=22$, 10 prices from \$50 to \$230, and a sweep over 160 values of $\eta\in[0,8000]$, the best choice $\eta=2167$ achieved normalized expected revenue of $78.0\%$ versus $71.0\%$ for RMS, an absolute improvement of $7.0\%$ [2203.11065]. The same paper reports that the MSE of $\phi$ estimation decreases monotonically with $\eta$ up to a sweet spot, after which over-exploration reduces revenue.

Revenue-maximizing ranking under random attention spans supplies another example of approximation-oriented learning curves. For random span $X$ with tail $G_x=\Pr(X\ge x)$, the Best-$x$ policy selects
$$
x^*\in\arg\max_x R(\sigma^x;x)G_x,
$$
where $\sigma^x$ is the fixed-span optimum [2012.03800]. Under the IFR condition on attention spans, Best-$x$ achieves at least $1/e$ of the clairvoyant benchmark, while no algorithm can exceed $1/2$ of that benchmark in the worst case [2012.03800]. In the contextual online version, RankUCB attains expected regret of order $\tilde O(\sqrt T)$ relative to the scaled $1/e$ benchmark despite censoring of the attention span [2012.03800].

The move from posted prices to richer mechanisms preserves the learning-curve perspective but increases parameter dependence. For menus of two-part tariffs, adversarial full-information online regret is
$$
\tilde O\!\big(\ell (K+H\ln H)\sqrt T\big),
$$
while fixed-length lottery menus admit
$$
\tilde O(m^2H\ell\sqrt T)
$$
full-information regret and
$$
\tilde O\!\big(m^2H\ell\,T^{1-1/(2\ell m+4)}\ln^{\ell m+1}(mHT)\big)
$$
bandit regret [2302.11700]. In the distributional setting, fixed-length lottery menus have sample complexity
$$
\tilde O\!\left(\frac{m^2H^2}{\epsilon^2}(\ell m+\ln(2/\delta))\right),
$$
whereas arbitrary-length menus exhibit the expected exponential dependence on $m$ under correlated valuations [2302.11700]. The same paper shows that dispersion-based methods, successful for smoothed online learning of two-part tariffs, are inadequate for menus of lotteries because dispersion can fail at the optimizer [2302.11700].

The open problems are correspondingly structural. One line asks when fixed-distribution curves can be sharpened beyond current worst-case universal rates, especially for unbounded supports and for ERM outside bounded-support or finite-support settings [2604.26922]. Another asks whether dimension-free final-phase guarantees survive multiple resource constraints in primal-dual inventory control [1812.09234]. In patient-buyer models, the gap between upper and lower bounds for mixed strategies and the possibility of $\sqrt T$ regret against optimal mixed menus remain unresolved [2202.06143]. In episodic time-varying revenue management, a formal regret analysis for dynamic per-period posterior sampling and a theory that quantifies the gains from correlated priors remain open [2405.04910].

Across these domains, the central lesson is not that revenue learning has a single canonical asymptotic law, but that it decomposes into sharply different regimes. Heavy tails without attainment can force arbitrarily slow convergence; finite optimal prices produce essentially $1/\sqrt n$ curves; discrete supports can yield almost exponential or exponential decay; nonstationarity introduces explicit variation exponents; and resource, feedback, or strategic constraints add separate structural terms that often dominate the statistical component.

Source: https://www.emergentmind.com/topics/learning-curves-for-revenue-maximization