Papers
Topics
Authors
Recent
Search
2000 character limit reached

Average-Cost Optimality Equation (ACOE)

Updated 23 February 2026
  • ACOE is a fundamental condition in Markov decision processes that defines optimality by balancing immediate costs with future value through a dynamic programming equation.
  • It is derived from discounted-cost models by taking the limit of relative value functions under conditions like weak continuity and inf-compact cost criteria.
  • ACOE underpins applications in queueing, inventory, and mean-field systems, enabling the computation of stationary deterministic optimal policies.

The average-cost optimality equation (ACOE) is a central object in the theory of Markov decision processes (MDPs) and stochastic control, providing a necessary and sufficient condition for a policy to be optimal with respect to long-run average (per-stage) cost. The ACOE connects dynamic programming with ergodic control, is foundational for the structure and computation of optimal policies, and provides the backbone for applications across queueing, inventory control, mean-field systems, and beyond.

1. Formal Statement of the ACOE

Let XX be a Borel subset of a Polish space (state space), A(x)AA(x) \subset A a Borel-action set (possibly noncompact) for each xXx \in X, c:X×A[0,]c: X \times A \to [0,\infty] a one-step cost, and P(x,a)P(\cdot\,|\,x,a) a transition probability kernel on XX. The average-cost optimality equation is

λ+h(x)=minaA(x){c(x,a)+Xh(y)P(dyx,a)},xX\lambda + h(x) = \min_{a \in A(x)} \left\{ c(x,a) + \int_X h(y)\,P(dy|x,a) \right\}, \quad \forall x\in X

where λR\lambda \in \mathbb R is the optimal average cost per unit time, and h:XRh:X \to \mathbb R is a "differential" or "relative" value function. The pair (λ,h)(\lambda,h) solves the ACOE under the following properties:

  • A(x)AA(x) \subset A0 is measurable (typically lower-semicontinuous),
  • A(x)AA(x) \subset A1 is inf-compact or A(x)AA(x) \subset A2 is A(x)AA(x) \subset A3-inf-compact over A(x)AA(x) \subset A4,
  • The minimum is attained, so measurable selectors A(x)AA(x) \subset A5 exist,
  • The integral A(x)AA(x) \subset A6 is finite when A(x)AA(x) \subset A7.

The ACOE governs the structure of stationary deterministic optimal policies and provides the threshold between Bellman-type optimality inequalities and actual equations. Solutions A(x)AA(x) \subset A8 are unique up to an additive constant in A(x)AA(x) \subset A9 (Feinberg et al., 2024).

2. Sufficient Conditions and Modern Existence Results

Contemporary results systematically weaken classical requirements on compactness, continuity, and uniform integrability. Key conditions ensuring the validity of the ACOE (see especially (Feinberg et al., 2024, Feinberg et al., 2012, Feinberg et al., 2016, Feinberg et al., 2017)) are:

  • Continuity/Compactness:
    • Weak model (W*): xXx \in X0 is xXx \in X1-inf-compact on xXx \in X2, xXx \in X3 is weakly continuous in xXx \in X4.
    • Setwise (S*): For each xXx \in X5, xXx \in X6 is inf-compact; xXx \in X7 is setwise continuous for all xXx \in X8.
  • Relative Value Boundedness:
    • c:X×A[0,]c: X \times A \to [0,\infty]0 for the discounted value.
    • c:X×A[0,]c: X \times A \to [0,\infty]1.
    • c:X×A[0,]c: X \times A \to [0,\infty]2.
    • Boundedness condition c:X×A[0,]c: X \times A \to [0,\infty]3: c:X×A[0,]c: X \times A \to [0,\infty]4 for each c:X×A[0,]c: X \times A \to [0,\infty]5.
    • Stronger c:X×A[0,]c: X \times A \to [0,\infty]6: c:X×A[0,]c: X \times A \to [0,\infty]7.
  • Equicontinuity/Integrability:
    • (EC) Uniform equicontinuity and a uniform integrable envelope for c:X×A[0,]c: X \times A \to [0,\infty]8.
    • (LEC) Lower-semi-equicontinuity in c:X×A[0,]c: X \times A \to [0,\infty]9, pointwise limit existence of P(x,a)P(\cdot\,|\,x,a)0, and uniform integrability w.r.t. P(x,a)P(\cdot\,|\,x,a)1 in P(x,a)P(\cdot\,|\,x,a)2.

Under W* or S*, P(x,a)P(\cdot\,|\,x,a)3, and LEC, the ACOE is satisfied; P(x,a)P(\cdot\,|\,x,a)4 with P(x,a)P(\cdot\,|\,x,a)5 is measurable, and policies selecting minimizers solve the average-cost control problem (Feinberg et al., 2024).

3. Derivation from Discounted to Average Cost, and Proof Techniques

The transition from the discounted-cost optimality equation to the ACOE is critical:

  • The value function P(x,a)P(\cdot\,|\,x,a)6 for P(x,a)P(\cdot\,|\,x,a)7 solves

P(x,a)P(\cdot\,|\,x,a)8

  • Define P(x,a)P(\cdot\,|\,x,a)9, XX0.
  • Under the boundedness assumptions, XX1 is pointwise bounded; diagonal/lower-semicontinuity arguments yield XX2.
  • Weakly continuous/inf-compact conditions permit passage of XX3 through XX4 and XX5, so the limiting function satisfies

XX6

Alternate approaches, such as occupation measure convex-analytic methods (employing ergodic occupation measures), Poisson/relative value iteration (RVI) schemes, and reduction to discounted MDPs (e.g., HV–AG transformation), also appear as foundational derivations (Arapostathis et al., 2019, Feinberg et al., 2015, Feinberg et al., 2017). The vanishing discount approach remains the standard, but split-chain constructions or Lyapunov drift stability hypotheses allow further generalization.

4. Policy Structure and Uniqueness

The ACOE under stated conditions admits solutions where for each XX7,

XX8

A measurable selector XX9 choosing a minimizer at each λ+h(x)=minaA(x){c(x,a)+Xh(y)P(dyx,a)},xX\lambda + h(x) = \min_{a \in A(x)} \left\{ c(x,a) + \int_X h(y)\,P(dy|x,a) \right\}, \quad \forall x\in X0 defines a deterministic stationary policy that is average-cost optimal. Any such policy solves both the average-cost optimality inequality and the equality, and achieves λ+h(x)=minaA(x){c(x,a)+Xh(y)P(dyx,a)},xX\lambda + h(x) = \min_{a \in A(x)} \left\{ c(x,a) + \int_X h(y)\,P(dy|x,a) \right\}, \quad \forall x\in X1 from every initial state.

The solution λ+h(x)=minaA(x){c(x,a)+Xh(y)P(dyx,a)},xX\lambda + h(x) = \min_{a \in A(x)} \left\{ c(x,a) + \int_X h(y)\,P(dy|x,a) \right\}, \quad \forall x\in X2 to the ACOE is unique up to a constant shift in λ+h(x)=minaA(x){c(x,a)+Xh(y)P(dyx,a)},xX\lambda + h(x) = \min_{a \in A(x)} \left\{ c(x,a) + \int_X h(y)\,P(dy|x,a) \right\}, \quad \forall x\in X3; i.e., if λ+h(x)=minaA(x){c(x,a)+Xh(y)P(dyx,a)},xX\lambda + h(x) = \min_{a \in A(x)} \left\{ c(x,a) + \int_X h(y)\,P(dy|x,a) \right\}, \quad \forall x\in X4 and λ+h(x)=minaA(x){c(x,a)+Xh(y)P(dyx,a)},xX\lambda + h(x) = \min_{a \in A(x)} \left\{ c(x,a) + \int_X h(y)\,P(dy|x,a) \right\}, \quad \forall x\in X5 both solve the ACOE and λ+h(x)=minaA(x){c(x,a)+Xh(y)P(dyx,a)},xX\lambda + h(x) = \min_{a \in A(x)} \left\{ c(x,a) + \int_X h(y)\,P(dy|x,a) \right\}, \quad \forall x\in X6 are bounded below (e.g., lower semicontinuous), then λ+h(x)=minaA(x){c(x,a)+Xh(y)P(dyx,a)},xX\lambda + h(x) = \min_{a \in A(x)} \left\{ c(x,a) + \int_X h(y)\,P(dy|x,a) \right\}, \quad \forall x\in X7 and λ+h(x)=minaA(x){c(x,a)+Xh(y)P(dyx,a)},xX\lambda + h(x) = \min_{a \in A(x)} \left\{ c(x,a) + \int_X h(y)\,P(dy|x,a) \right\}, \quad \forall x\in X8. This is a generalization of the classical uniqueness theorem for the Bellman equation in ergodic control (Feinberg et al., 2024).

5. Comparison with Classical and Alternative Conditions

Classically, average-cost optimality analysis relied on:

  • Communicating or unichain structure for finite-state models,
  • Lyapunov (drift) conditions to ensure positive recurrence,
  • Uniform compactness of action sets, and strong Feller continuity of transitions,
  • Uniform equicontinuity of discounted value functions.

Recent advances weaken these requirements, replacing them by:

  • λ+h(x)=minaA(x){c(x,a)+Xh(y)P(dyx,a)},xX\lambda + h(x) = \min_{a \in A(x)} \left\{ c(x,a) + \int_X h(y)\,P(dy|x,a) \right\}, \quad \forall x\in X9-inf-compactness or inf-compactness of costs rather than action set compactness,
  • Weak or setwise continuity instead of strong-Feller continuity,
  • One-sided boundedness in the limit of discounted relative values,
  • Lower-semi-equicontinuity and uniform integrability (LEC) instead of full equicontinuity.

This encompasses a broader range of stochastic control models, such as queueing or inventory systems with noncompact action sets and weak continuity properties, which often fall outside the reach of classical assumptions (Feinberg et al., 2024, Feinberg et al., 2016, Feinberg et al., 2012).

6. Illustrative Examples and Applications

The wide applicability of the ACOE is demonstrated by explicit examples in (Feinberg et al., 2024):

  • Single-Action Indicator–Cost: λR\lambda \in \mathbb R0, λR\lambda \in \mathbb R1, λR\lambda \in \mathbb R2, λR\lambda \in \mathbb R3. The relative value function λR\lambda \in \mathbb R4 is lower-semicontinuous, and λR\lambda \in \mathbb R5.
  • Dirichlet–Cost MDP: λR\lambda \in \mathbb R6, λR\lambda \in \mathbb R7, λR\lambda \in \mathbb R8, λR\lambda \in \mathbb R9 (Dirichlet function). The bias function h:XRh:X \to \mathbb R0 fails to be lower-semicontinuous but the ACOE still holds under the weaker integrability and limit assumptions.

Broader applications include the derivation of optimal h:XRh:X \to \mathbb R1 policies in inventory systems, mean-field game limit problems, and models with state- or action-dependent control constraints. Computationally, the ACOE provides the foundation for value iteration, policy iteration, and linear programming methods for average-cost control (Feinberg et al., 2024, Feinberg et al., 2016, Arapostathis et al., 2019).

7. Impact and Extensions

The ACOE is essential for the theoretical and computational treatment of Markov control problems under the average-cost criterion. Its solution structure and existence theory underpin modern direct algorithms and facilitate the analysis of ergodic control for noncompact, weakly continuous, and complex stochastic dynamic models. The recent generalizations to weaker boundedness and continuity, as established in (Feinberg et al., 2024), have expanded its reach to previously intractable classes of queueing, inventory, and stochastic network models. Furthermore, its connections with ergodic occupation measures, split-chain Poisson equations, and mean-field limits continue to drive advances in both theory and large-system applications.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Average-Cost Optimality Equation (ACOE).