---
title: Least Cost Principle in Optimization
url: https://www.emergentmind.com/topics/least-cost-principle
type: topic
---

# Least Cost Principle in Optimization

Searching arXiv for the cited papers to ground the article in the current record.
arxiv_search(query="1112.2397 OR \"Optimal posting price of limit orders: learning by trading\" OR \"The Cost of Optimally Acquired Information\" OR 2511.05466 OR \"Optimal pricing for optimal transport\" OR 1405.4809", max_results=10)
Least Cost Principle appears in several distinct technical senses. In algorithmic trading, it prescribes moving the posting distance in the direction that reduces expected marginal cost; in information acquisition, indirect costs arise as minimal expected costs under flexible sequential acquisition; in optimal transport, it is the minimization of total transport cost over feasible transport plans; and in a continuous-time control formulation of physics, it is the minimization of a discounted integral of an acceleration cost minus a state-dependent reward [1112.2397] [2511.05466] [1405.4809] [2603.25444]. This suggests a common optimization motif: a feasible object is selected by minimizing a structurally constrained cost functional, with duality, convexity, or dynamic-consistency conditions determining existence, uniqueness, and implementability.

## 1. Scope of the term and formal pattern

In the cited literature, the optimization variable can be a query plan \(p\), a posting distance \(\delta\), a random posterior \(\pi\), a transport plan \(\pi\in\Pi(\mu,\nu)\), or an acceleration path \(a(\cdot)\). The associated objective can be a scalar cost model \(C(p,\theta)\), a penalized expected execution cost \(C(\delta)\), an indirect information cost \(\Phi(C)(\pi)\), a transport cost \(\int c\,d\pi\), or a discounted cost-to-go \(\mathcal{C}[x(\cdot),v(\cdot),a(\cdot)]\).

A database formulation makes the contrast especially explicit: a System R–style optimizer typically chooses the plan of least cost given some fixed value of the parameters, whereas least expected cost query optimization chooses the plan of the least expected cost and does not rely on the assumptions that it is enough to optimize for the expected case or that the parameters are constant throughout the execution of the query [9909016].

The formal pattern is therefore not tied to a single domain-specific object. What recurs is the replacement of heuristic choice by an optimization problem over admissible objects, together with structural conditions that make the minimizer meaningful: convexity in limit-order placement, sequential learning-proofness in information acquisition, Kantorovich duality and \(c\)-convexity in optimal transport, and Hamilton–Jacobi–Bellman optimality in the continuous-time control formulation of physics.

## 2. Limit-order execution as a least-cost posting problem

In continuous auctions, a trader operates over short, repeated posting periods of length \(T\), sending one passive order of size \(Q_T\) at the beginning of each period at a distance \(\delta\) from a reference fair price process \((S_t)_{t\in[0,T]}\). At the end of the period, any unexecuted remainder is immediately completed with a market order and incurs a market-impact penalty. The execution flow of a buy order posted at price \(S_0-\delta\) is modeled as a Poisson process \(N^{(\delta)}\) with random intensity
\[
\Lambda_T(\delta,S) := \int_0^T \lambda(S_t-(S_0-\delta))\,dt,
\]
where \(\lambda:[-S_0,+\infty)\to\mathbb{R}_+\) is finite, non-increasing, and convex. The realized expected cost is
\[
C(\delta):=E\Big[(S_0-\delta)(Q_T\wedge N_T^{(\delta)})+\kappa S_T\Phi((Q_T-N_T^{(\delta)})_+)\Big],
\]
and the objective is to minimize \(C(\delta)\) over \([0,\delta_{\max}]\) [1112.2397].

The paper derives an expected gradient representation \(C'(\delta)=E[H(\delta,S)]\) and implements a projected stochastic approximation,
\[
\delta_{n+1}=\mathrm{Proj}_{[0,\delta_{\max}]}\Big(\delta_n-\gamma_{n+1}H(\delta_n,S^{(n+1)})\Big),
\]
with \(H(\delta_n,S^{(n+1)})\) computed from the observed path through \(\Lambda_T\) and its \(\delta\)-derivatives. In practice the path can be discretized, yielding the implementable update
\[
\delta_{n+1}=\mathrm{Proj}_{[0,\delta_{\max}]}\Big(\delta_n-\gamma_{n+1}H(\delta_n,(\bar S^{(n+1)}_{t_i})_{0\le i\le m})\Big).
\]
Under standard step-size conditions, moment bounds, and strict monotonicity of the mean field \(h(\delta)=E[H(\delta,S)]\), the projected stochastic gradient converges almost surely to a unique least-cost posting distance \(\delta^*\in(0,\delta_{\max})\) [1112.2397].

A central structural tool is the functional co-monotony principle for one-dimensional diffusions. It yields verifiable sufficient conditions, stated in terms of model parameters and simple functionals of \(S\), ensuring \(C'(0)<0\) and \(C''(\delta)\ge 0\) on \([0,\delta_{\max}]\). In the exponential intensity case \(\lambda(x)=Ae^{-kx}\), these conditions become explicit inequalities involving \(k\), \(\kappa\), \(Q_T\), and \(E[S_T]\). The numerical experiments reported in the paper show that the cost \(C(\delta)\) is strictly convex with a unique minimum, and that one stochastic-approximation run converges substantially faster than brute-force Monte Carlo evaluation of the full curve [1112.2397].

## 3. Information acquisition and sequential minimization

In the information-acquisition framework, the decision-maker faces a finite state space \(\Theta\), beliefs are probability vectors in \(\Delta(\Theta)\), and an experiment induces a random posterior \(\pi\in R:=\Delta(\Delta(\Theta))\). A direct cost is any map \(C:R\to\overline{\mathbb{R}}_+\) such that \(C[R_{\emptyset}]=\{0\}\). The indirect cost is generated by sequential minimization. For a two-step policy \(\Pi\), the two-step learning map is
\[
\Psi(C)(\pi):=\inf_{\Pi\in\Delta^\dagger(R)}\Big\{C(\pi_1)+E_\Pi[C(\pi_2)]\Big\}
\quad\text{subject to}\quad E_\Pi[\pi_2]\ge_{mps}\pi,
\]
and the sequential learning map is
\[
\Phi(C)(\pi):=\lim_{n\to\infty}\Psi^n(C)(\pi).
\]
The envelope characterization states that \(\Phi(C)\) is the largest sequential-learning-proof cost below the direct cost [2511.05466].

Sequential learning-proofness (SLP) is the recursive fixed-point property \(\Psi(C)=C\). The characterization theorem states that \(C\) is an indirect cost if and only if \(C\) is SLP, and that this is equivalent to \(C\) being Monotone and Subadditive. Monotonicity is defined by \(C(\pi)\le C(\pi')\) whenever \(\pi\le_{mps}\pi'\), while subadditivity requires
\[
C(E_\Pi[\pi_2])\le C(\pi_1)+E_\Pi[C(\pi_2)]
\]
for all finite-support two-step policies. In economic terms, SLP rules out “cost arbitrage” through sequential decomposition [2511.05466].

A major class of solutions is uniformly posterior separable (UPS) costs. These take the form
\[
C^{H}_{ups}(\pi)=E_\pi\big[H(q)-H(p_\pi)\big]
\]
for a convex potential \(H\). Under the paper’s Regularity condition, SLP and Regularity are equivalent to the existence of such a \(C^1\) convex potential on an open convex domain. Mutual information and Wald/MS costs are presented as canonical examples, and the paper introduces two additional indirect cost functions: Total Information (TI), which is UPS and Additive, and Minimal Likelihood Ratio (MLR), which is SLP and Prior Invariant but not Regular/UPS [2511.05466].

The framework also sharpens a central controversy in rational inattention: the information cost trilemma. For any nontrivial cost with rich domain, SLP and Constant Marginal Cost are equivalent to a Total Information cost; Prior Invariance, Constant Marginal Cost, and Dilution Linearity characterize an LLR cost; and if a cost is the rich-domain restriction of MLR, then it is SLP and Prior Invariant. Conversely, SLP plus Prior Invariance implies not Constant Marginal Cost and not UPS. The paper’s conclusion is that modelers cannot have all three of SLP, PI, and CMC simultaneously for nonzero costs [2511.05466].

## 4. Optimal transport, dual prices, and constrained pricing envelopes

In optimal transport, the least-cost formulation is the Monge–Kantorovich primal problem. Given probability measures \(\mu\) on \(X\) and \(\nu\) on \(Y\), and a transport cost \(c:X\times Y\to\mathbb{R}\), the admissible set is \(\Pi(\mu,\nu)\), the probability measures on \(X\times Y\) with marginals \(\mu\) and \(\nu\). The primal problem is
\[
\inf_{\pi\in\Pi(\mu,\nu)}\int_{X\times Y} c(x,y)\,d\pi(x,y).
\]
This is the least cost principle in its most literal form: among all feasible routings of mass from \(\mu\) to \(\nu\), select a plan minimizing total transport cost [1405.4809].

Kantorovich duality supplies the pricing interpretation. The dual problem is
\[
\sup_{f,g}\left\{\int_X f(x)\,d\mu(x)+\int_Y g(y)\,d\nu(y): g(y)-f(x)\le c(x,y)\ \forall x,y\right\}.
\]
If \((f,g)\) are optimal and \(\pi\) is an optimal plan, then
\[
g(y)-f(x)=c(x,y)
\quad\text{for }\pi\text{-almost every }(x,y).
\]
Thus \(f(x)\) acts as a source price, \(g(y)\) as a destination price, the inequality encodes feasibility of the markup, and equality holds on transported pairs. The paper formulates a constrained optimal pricing problem in which part of the optimal transportation plan is kept fixed and some source prices are also fixed, and solves it using \(c\)-convexity, \(c\)-transforms, and \(c\)-antiderivatives [1405.4809].

Given a mapping \(M:X\rightrightarrows Y\), a \(c\)-antiderivative \(f\), and a nonempty set \(S\subset\mathrm{dom}(M)\), the admissible family is
\[
\mathcal{A}[c,f|_S,M]
:=\Big\{h:X\to(-\infty,+\infty]\ \text{\(c\)-convex}:\
G(M)\subset G(\partial_c h),\ h|_S=f|_S\Big\}.
\]
The main existence theorem states that this family is nonempty and contains a lower envelope
\[
\underline h(x)=\inf\{h(x):h\in\mathcal{A}[c,f|_S,M]\}=:\mathsf{@}[c,f|_S,M](x)
\]
and an upper envelope
\[
\overline h(x)=\sup\{h(x):h\in\mathcal{A}[c,f|_S,M]\}=:\mathsf{Y}[c,f|_S,M](x).
\]
The lower envelope has the explicit Rockafellar-type form
\[
\mathsf{@}[c,f|_S,M](x)=\sup_{s\in S}\big[f(s)+R[c,M,s](x)\big],
\]
where \(R[c,M,s]\) is defined by chains in the graph of \(M\). In the metric case \(X=Y=\mathbb{R}^n\), \(c(x,y)=\|x-y\|\), and \(M=\mathrm{Id}\), these envelopes reduce to the McShane–Whitney formulas for the lowest and highest 1-Lipschitz extensions [1405.4809].

## 5. Discounted acceleration cost and inverse-optimal physics

A recent control-theoretic formulation states the Least Cost Principle as an infinite-horizon optimal control problem for \(N\) particles with positions \(x_i(t)\), velocities \(v_i(t)=\dot x_i(t)\), and accelerations \(a_i(t)=\dot v_i(t)\). The discounted cost functional is
\[
\mathcal{C}[x(\cdot),v(\cdot),a(\cdot)]
=
\int_{t_0}^{\infty} e^{-\gamma (t-t_0)}
\left(
\sum_{i=1}^{N}\frac{1}{2}m_i\|a_i(t)\|^2
-
R(x(t),v(t))
\right)\,dt,
\]
with fixed initial conditions and dynamics \(\dot x_i=v_i\), \(\dot v_i=a_i\). The principle asserts that the realized physical evolution is the solution to
\[
\mathcal{C}^*(x_0,v_0)=\inf_{a(\cdot)}\mathcal{C}[x(\cdot),v(\cdot),a(\cdot)].
\]
The paper derives the quadratic acceleration cost from time homogeneity, spatial isotropy, additivity over particles and masses, and invariance across homogeneously accelerated frames [2603.25444].

The optimality condition is
\[
m_i\ddot x_i=-\,\frac{\partial \mathcal{C}^*(x,v)}{\partial v_i}.
\]
If the observed dynamics obey Newton’s second law \(m_i\ddot x_i=F_i(x)\), then the optimal cost-to-go can be written as
\[
\mathcal{C}^*(x,v)=-\sum_{i=1}^{N}F_i(x)\cdot v_i + B(x),
\]
where \(B(x)\) is any function of position only. Substitution into the Hamilton–Jacobi–Bellman identity yields the inverse-optimal reward
\[
R(x,v)
=
-\sum_{i,j} v_i^\top (\nabla_{x_j}F_i(x))\,v_j
-\frac{1}{2}\sum_{i=1}^N \frac{\|F_i(x)\|^2}{m_i}
+\sum_{i=1}^N v_i\cdot\big(\gamma F_i(x)+\nabla_{x_i}B(x)\big)
-\gamma B(x).
\]
The mapping \(F\mapsto R\) is therefore unique only up to the addition of \(B(x)\) and its induced linear-in-velocity term [2603.25444].

For Newtonian gravitation, the reward decomposes into a positive term proportional to relative speed squared, a negative term penalizing radial motion, and a force-coupling term quadratic in \(G\). In the two-body case, the contribution of the first two terms reduces to
\[
R_{ij}^{(\mathrm{I+II})}
=
\frac{Gm_im_j}{2r_{ij}^3}\big(\|u_\perp\|^2-2u_\parallel^2\big),
\]
so for fixed \(\|u_{ij}\|\) and \(r_{ij}\), the reward is maximized by \(u_\parallel=0\), that is, purely tangential relative motion. The paper therefore interprets the inferred reward as favoring tangential trajectories and quasi-circular motion at short separations. For Coulomb forces, the same algebraic structure holds after replacing \(Gm_im_j\) by \(-C_{\mathrm{Coulomb}}q_iq_j\), so attraction and repulsion reverse the reward effects of speed and tangentiality according to the sign of \(q_iq_j\) [2603.25444].

## 6. Structural properties, limitations, and cross-domain interpretation

Several structural themes recur. In the trading formulation, almost sure convergence requires strict convexity of \(C\), sign conditions such as \(C'(0)<0\), standard step-size summability, and moment bounds; in the real-data averaging case it additionally requires a pathwise Lyapunov monotonicity and discrepancy conditions for the averaged sequence [1112.2397]. In the information framework, the central structural property is SLP, which is equivalent to monotonicity and subadditivity, while kernel bounds show that local curvature cannot be reduced by optimization [2511.05466]. In optimal transport, feasible prices are organized by \(c\)-convex envelopes, and complementary slackness identifies the equality set on the support of an optimal plan [1405.4809]. In the physics formulation, the reward is recovered only up to \(B(x)\), and the discount factor \(\gamma>0\) is introduced for tractability and to ensure convergence of the infinite-horizon integral [2603.25444].

The limitations are equally domain-specific. The trading model uses a non-homogeneous Poisson execution flow whose intensity depends only on distance to \(S_t\), and it does not model queue dynamics, hidden liquidity, or simultaneous multi-level orders [1112.2397]. The information-acquisition framework assumes finite \(\Theta\), Polish signal spaces, and full flexibility over sequential policies, with no discounting or time preference in the baseline reduced form [2511.05466]. The optimal-transport pricing theory establishes existence of extremal constrained prices in general spaces with lower semicontinuous costs, but the paper explicitly notes that uniqueness of the constrained family is a natural question left for future study [1405.4809]. The physics formulation assumes central or position-dependent forces, rewards depending on \(x\) and \(v\) but not on \(a\), and smoothness and boundedness conditions sufficient for the HJB derivations [2603.25444].

A common misconception is to treat “least cost” as a single theorem or a universally identical variational principle. The cited literature does not support that reading. One database usage contrasts least expected cost with optimization at fixed parameter values [9909016]; one market-microstructure usage builds a projected stochastic gradient for a unique least-cost posting distance [1112.2397]; one information-theoretic usage defines indirect cost as the fixed point of a sequential minimization operator [2511.05466]; one optimal-transport usage couples transport-cost minimization with extremal compatible pricing policies [1405.4809]; and one physical usage reinterprets laws of motion as the minimizers of a discounted acceleration-cost functional [2603.25444]. This suggests that “Least Cost Principle” is best understood as a family of optimization doctrines unified by constrained minimization, but differentiated by the structure of admissible objects, the meaning of cost, and the specific regularity conditions that make minimization analytically and computationally tractable.

Source: https://www.emergentmind.com/topics/least-cost-principle