---
title: Utility–Information Equivalence
url: https://www.emergentmind.com/topics/utility-information-equivalence
type: topic
---

# Utility–Information Equivalence

Utility–Information Equivalence denotes a family of results in which utility, reward, or value functions are placed in exact correspondence with probability, entropy, relative entropy, or information-processing costs. In the bounded-rationality line developed by Ortega and Braun, the central statement is that under simple choice axioms, utilities and log-probabilities are linearly related, so optimal bounded-rational behavior takes a Gibbs form and is selected by a variational free-utility principle [1107.5766]. Closely related work derives reward from negative information content, identifies expected utility with negative entropy rate or negative cross-entropy in stochastic processes and agent–environment coupling, and extends the same bridge to rational inattention, value of information, privacy, information leakage, and betting with side information [0911.5106] [1709.09117] [2510.11102] [1102.3751] [2306.07975].

## 1. Axiomatic conversion between utility and probability

A canonical formulation starts from a finite probability space \((\Omega,\mathcal F,P)\) and models decision-making by a probability distribution over elementary events \(\omega\in\Omega\). The decision maker is assumed to admit a real-valued utility function \(U\), with incremental form
\[
u(A\mid B):=U(A\cap B)-U(B),
\]
satisfying three desiderata: real-valuedness, additivity, and monotonicity. In the formulation of "Information, Utility & Bounded Rationality," these are stated as \(u(A\mid B)=f(P(A\mid B))\), \(u(A\cap B\mid C)=u(A\mid C)+u(B\mid A\cap C)\), and
\[
P(A\mid B)>P(C\mid D)\iff u(A\mid B)>u(C\mid D).
\]
A celebrated consequence is that \(f\) must be logarithmic, so there exists a conversion constant \(\alpha>0\) such that
\[
U(A\cap B)-U(B)=\alpha\log P(A\mid B)
\]
for all events \(A,B\) [1107.5766].

For singletons \(\omega\), the same relation implies that the chosen distribution is a Gibbs measure,
\[
P(\omega)=\frac{\exp[(1/\alpha)U(\omega)]}{Z},
\qquad
Z=\sum_\omega \exp[(1/\alpha)U(\omega)].
\]
The 2010 axiomatic formalization presents the same conclusion with an additive constant,
\[
U(A\mid B)=\alpha\log P(A\mid B)+\beta,
\]
emphasizing that only utility differences matter [1007.0940].

A closely related 2009 treatment formulates the same logic in terms of rewards. Starting from real-valuedness, additivity, and order-preservation, rewards are uniquely determined by negative information content. Extending from single events to stochastic processes, the utility of a realization is defined as its reward rate, and the expected utility of a stochastic process becomes the negative entropy rate. In coupled agent–environment systems, the expected utility actually achieved by the agent is the negative cross-entropy from the input-output distribution of the coupled interaction system and the agent’s input-output distribution [0911.5106].

These derivations establish the strongest version of utility–information equivalence: utility is not merely regularized by information, but is induced by probability itself through a logarithmic conversion law.

## 2. Free utility and KL-regularized bounded rationality

Once utility and probability are conjugate, the theory can be written as a variational principle. The free-utility functional is
\[
J[P;U]=\sum_\omega P(\omega)\,U(\omega)-\alpha\sum_\omega P(\omega)\log P(\omega).
\]
By Gibbs’ inequality, the Gibbs law uniquely maximizes \(J[P;U]\), and the optimal value is \(U(\Omega)\) [1107.5766].

The most widely used bounded-rational form arises when a decision maker begins from a default policy \(p_0(\omega)\) and faces an added payoff \(U(\omega)\). Renaming the inverse temperature as \(\beta=1/\alpha\), the decision problem becomes
\[
q^*=\arg\max_q\left\{\sum_\omega q(\omega)\,U(\omega)-\frac{1}{\beta}D_{KL}[q\|p_0]\right\}.
\]
Its stationary condition yields
\[
q^*(\omega)=\frac{p_0(\omega)\exp[\beta U(\omega)]}{Z(\beta)},
\qquad
Z(\beta)=\sum_\omega p_0(\omega)\exp[\beta U(\omega)].
\]
The optimization therefore trades off expected utility gain against an information-processing cost proportional to \(D_{KL}(q\|p_0)\) [1107.5766].

The 2010 formalization presents the same transformation as a change in free utility caused by the addition of new constraints expressed by a target utility function. In that presentation, the new policy maximizes expected gain minus the information cost of deviating from a reference distribution, and the bounded-rational policy takes the explicit form
\[
P_f(x)=
\frac{P_0(x)\exp\bigl(\tfrac{1}{\alpha}U_*(x)\bigr)}
{\sum_{x'}P_0(x')\exp\bigl(\tfrac{1}{\alpha}U_*(x')\bigr)}.
\]
This is the same Gibbs reweighting, written relative to a prior policy rather than an absolute utility scale [1007.0940].

The parameter \(\beta\) quantifies computational or information-processing resources. As \(\beta\to 0\), the policy converges to the default \(p_0\); as \(\beta\to\infty\), it collapses onto \(\arg\max_\omega U(\omega)\). The classical maximum-expected-utility principle is therefore recovered when resource costs are ignored [1107.5766].

## 3. Thermodynamic reading and sequential control

The bounded-rationality framework is explicitly thermodynamic in form. The correspondences given in the 2011 paper are: energy with negative utility, entropy with the uncertainty penalty in \(J\), free energy with free utility, and work required to change a distribution \(p_0\to q\) with
\[
W=\alpha D_{KL}(q\|p_0).
\]
In the physical analogy, any information update that reshapes probabilities carries a thermodynamic cost, and \(\alpha=kT\) maps bits to energy [1107.5766].

In sequential decision problems, the same variational principle is applied recursively to passive or default dynamics \(p_0(x_1), p_0(x_2\mid x_1),\ldots\) together with stagewise scalar payoffs. This yields nested Gibbs updates such as
\[
q(x_2\mid x_1)=\frac{p_0(x_2\mid x_1)\exp[\beta U(x_2\mid x_1)]}{Z_2(x_1)},
\]
\[
q(x_1)=\frac{p_0(x_1)\exp[\beta(U(x_1)+V_2(x_1))]}{Z_1},
\qquad
V_2(x_1)=\frac{1}{\beta}\log Z_2(x_1).
\]
Unless \(\beta\to\infty\), the resulting control law remains stochastic. The paper identifies this as the hallmark of resource-bounded optimal control: the policy naturally carries entropy to economize on search [1107.5766].

The same framework also subsumes risk-sensitive and robust control. If the environment is modeled as a bounded-rational opponent with finite \(\beta_{\mathrm{env}}\), its optimal response is again Gibbs with its own inverse temperature. Positive \(\beta_{\mathrm{env}}\) corresponds to risk-seeking, negative \(\beta_{\mathrm{env}}\) to risk-averse or pessimistic variational forms recovering the exponential-utility criterion, and the limit \(\beta_{\mathrm{env}}\to-\infty\) yields worst-case minimization. In that limit, the outer policy becomes
\[
\arg\max_x \min_y U(x,y),
\]
the zero-sum minimax solution [1107.5766].

The 2010 axiomatic paper states the same broader consequence in application terms: optimal control, adaptive estimation, and adaptive control problems can be solved in this way in a resource-efficient way, and the maximum expected utility rule is recovered when resource costs are ignored [1007.0940].

## 4. Equivalence theorems beyond bounded rationality

A distinct but closely related equivalence result links additive random-utility discrete choice to rational inattention. In the discrete-choice model, utility takes the form
\[
U_i=v_i+\varepsilon_i,
\]
and the ex-ante surplus is
\[
W(v)=E_\varepsilon\!\left[\max_i\{v_i+\varepsilon_i\}\right].
\]
Choice probabilities are the gradient of the surplus, \(q_i(v)=\partial W(v)/\partial v_i\). In the rational inattention formulation, information costs are defined by a generalized entropy function \(H_S\), and the decision maker solves
\[
\max_{p(v)\in\Delta}\left\{E_v[v\cdot p(v)]-I_S[p(\cdot)]\right\},
\]
where
\[
I_S[p(\cdot)]=H_S(E_v[p(v)])-E_v[H_S(p(v))].
\]
The main theorem states that the unique solution satisfies
\[
p_i(v)=q_i(v)=\frac{\partial W(v)}{\partial v_i}.
\]
Thus generalized-entropy rational inattention and additive random utility are observationally equivalent. The paper states that any additive random utility model can be given an interpretation in terms of boundedly rational behavior, including multinomial logit, nested logit, and probit [1709.09117].

Another line of work studies utility–information equivalence through value of information and separable utility. In a finite state space \(S\), a utility act is a vector \(u\in\mathbb R^n\), beliefs are \(p\in\Delta(S)\), and the expected utility pairing is
\[
\langle u,p\rangle=\sum_{s\in S}u(s)p(s).
\]
Decision makers are represented by continuous closed convex comprehensive c-utility-act sets. The central theorem proves the equivalence of three conditions: \(M\) values information more than \(L\); \(m-\ell\) is convex on \(\Delta(S)\); and there exists a c-utility set \(T\) such that
\[
M=L+T.
\]
Here \(+\) is Minkowski addition. Economically, \(X+Y\) corresponds to “fusing” two sets of decisions: one chooses \(a\in X\) and \(b\in Y\), obtaining payoff-vector \(a+b\). The same paper introduces a dioid structure under union \((\oplus)\) and fusion \((\otimes)\), and states
\[
M \text{ is VoI-higher than } L \iff M \text{ is fusion-more-flexible than } L.
\]
In this formulation, increasing value of information is equivalent to additive separability via decision fusion [2510.11102].

These results do not assert the same identity as the logarithmic conversion law, but they preserve the same structural theme: information value can be characterized exactly by a utility representation, and vice versa.

## 5. Privacy, leakage, information pricing, and betting

In database sanitization, utility–information equivalence is formulated operationally. With private variables \(X\), public variables \(Y\), and sanitized output \(Z\), utility is measured by information preserved about \(Y\),
\[
U(p(z\mid y))=I(Y;Z),
\]
while privacy is measured by leakage about \(X\),
\[
L(p(z\mid y))=I(X;Z)\le \varepsilon.
\]
Equivalently, one may impose a distortion constraint \(E[d(Y,Z)]\le D\). The optimal sanitization channel solves
\[
\max_{p(z\mid y)} I(Y;Z)
\quad\text{s.t.}\quad
I(X;Z)\le\varepsilon,\;\;E\,d(Y,Z)\le D.
\]
The paper states that in this framework utility is literally “the information retained about \(Y\),” privacy is “the information leaked about \(X\),” and the trade-off is solved by a single mutual-information optimization. Closed-form boundaries are given for categorical data with Hamming distortion via reverse-waterfilling and for jointly Gaussian data with MSE via additive Gaussian noise [1102.3751].

In information-leakage games, utility is also defined directly by information. The paper considers zero-sum attacker–defender games in which the payoff is either posterior vulnerability in Quantitative Information Flow or the max-divergence that defines \(\varepsilon\)-differential privacy. A key novelty is that the utility is given by information leakage, and this deviates from classical game theory because the payoff is not linear with respect to mixed strategies. Nevertheless, the framework establishes the existence of Nash equilibria for both QIF-games and DP-games and provides computational procedures based on projected subgradient descent or linear programming [2012.12060].

A financial formulation treats information as a quantity that converts a prior distribution into a posterior distribution, with amount measured by relative entropy
\[
D_{KL}(Q\|P).
\]
Pricing is based on a utility-indifference condition equating the maximized a posterior utility after paying cost \(C\) with the utility of the a priori optimal strategy. In the Gaussian–exponential example, the information cost is quadratic in the signal realization, and the paper reports that the cost grows linearly to first order with \(D_{KL}(Q\|P)\). It further distinguishes one price for a given quantity of upside information and another price for a given quantity of downside information [1106.5706].

Betting-theoretic work extends the equivalence to isoelastic utility of wealth and utility of wealth ratios. In betting tasks with side information, maximized isoelastic certainty-equivalent utility yields conditional Rényi divergences; in the two-parameter formulation, the mutual information \(I^{\rm ID}_{q,r}(X;G)\) is exactly the utility assigned to the ratio between best-with-side-information and best-without-side-information certainty-equivalent payoffs. The same quantities are then carried over to quantum state-betting and noisy quantum state-betting, where they operationally characterize measurement informativeness and non-constancy of channels [2306.07975].

## 6. Scope, distinctions, and common misconceptions

The phrase “Utility–Information Equivalence” does not designate a single universally identical theorem. The literature represented here contains at least four distinct but related meanings.

| Setting | Utility side | Information side |
|---|---|---|
| Bounded rationality | \(E_q[U]-\alpha D_{KL}(q\|p_0)\) | KL cost, Gibbs law |
| Stochastic processes | expected utility | negative entropy rate, negative cross-entropy |
| Rational inattention | surplus \(W\) or entropy \(H\) | generalized entropy cost |
| Privacy and leakage | \(I(Y;Z)\) or leakage utility | preserved information or leaked information |

First, in the axiomatic bounded-rationality papers, the equivalence is literal: under real-valuedness, additivity, and monotonicity, utility and log-probability are linearly related [1107.5766] [1007.0940]. Second, in the stochastic-process formulation, expected utility becomes a negative information measure such as entropy rate or cross-entropy [0911.5106]. Third, in rational inattention and discrete choice, the equivalence is convex-analytic and observational: different primitives generate the same choice probabilities [1709.09117]. Fourth, in privacy, leakage, and betting, utility is defined directly by an information-theoretic quantity or by an operational advantage due to side information [1102.3751] [2012.12060] [2306.07975].

A recurrent misconception is that introducing information into utility must eliminate stochasticity. The bounded-rational control result states the opposite: unless \(\beta\to\infty\), the optimal policy remains stochastic, and the entropy of the policy is part of the resource trade-off [1107.5766]. Another misconception is that the framework replaces classical expected-utility theory. In the Ortega–Braun formulation, maximum expected utility is recovered when resource costs are ignored; in the equivalent limit, the policy concentrates on the maximizer of utility [1107.5766].

A further point of clarification is that “information” can appear either as a cost or as an objective, depending on the problem. In bounded rationality it enters as an information-processing cost \(\alpha D_{KL}(q\|p_0)\); in sanitized databases it is utility preserved about public attributes and privacy leaked about private attributes; in information-leakage games it is the zero-sum payoff itself [1107.5766] [1102.3751] [2012.12060].

A related but distinct use of the theme appears in multi-agent utility design with partial information networks. There, numerical case studies show a practical trade-off between improving communication links and designing more tailored local utilities to optimize the Price of Anarchy. This suggests a broader interpretation in which information structure and utility design can sometimes substitute for one another at the level of system performance, even when no direct logarithmic conversion law is involved [2501.17385].

Across these formulations, the unifying idea is precise rather than metaphorical: information is not merely adjacent to utility. Depending on the formal setting, it can determine utility through a logarithmic conversion, regularize utility maximization through relative entropy, coincide with expected utility through entropy-based identities, or constitute the utility function itself.

Source: https://www.emergentmind.com/topics/utility-information-equivalence