---
title: 'Information Value (IV): Concepts & Applications'
url: https://www.emergentmind.com/topics/information-value-iv
type: topic
---

# Information Value (IV): Concepts & Applications

Information Value (IV) denotes a family of quantitative constructs that assign value to information under a specified model of uncertainty, utility, prediction, or discrimination. In decision theory, IV commonly refers to the increase in expected utility or the reduction in Bayes risk induced by information, often under explicit information constraints or experiment structures [1405.5860]. In hidden-state estimation and networked control, it appears either as mutual information between a current latent status and received observations, or as the decrement in a control-oriented value function caused by allowing a transmission [2008.08907]. In natural-language processing, it is defined as the distance between an utterance and a set of plausible alternatives [2310.13676]. In binary feature screening, it is the Jeffreys divergence between class-conditional bin distributions [2309.13183]. This multiplicity of meanings is not accidental: the term is used whenever information is evaluated relative to a task, a decision rule, or a predictive objective.

## 1. Major formalizations

Several distinct research traditions use the label “Information Value” or “Value of Information,” each with its own state space, action space, and optimization criterion.

| Domain | Formal object | Representative definition |
|---|---|---|
| Decision under uncertainty | Utility gain under information constraint | $\overline u_S(\lambda)=\sup_{P(B|A)}\{E_P[u(A,B)]:I_S(A;B)\le \lambda\}$ |
| Belief-based decision problems | Gain from posterior randomization | $\mathrm{VoI}_A(q)=E[V(q)]-V(p)$ |
| Hidden Markov or latent-variable models | Uncertainty reduction about current state | $v(t)=I(X_t;Y_{t_1'},\dots,Y_{t_n'})$ |
| Networked control | Difference in cost-to-go with and without transmission | $\mathrm{voi}_k=V^e_k|_{\delta_k=0}-V^e_k|_{\delta_k=1}$ |
| Dialogue and text comprehension | Distance from plausible alternative utterances | $I(Y=y\mid X=x)=d(y,A_x)$ |
| Binary feature selection | Symmetric divergence between class-conditionals | $\mathrm{IV}=J(\mathbf p,\mathbf q)=\sum_j(p_j-q_j)\ln\frac{p_j}{q_j}$ |

In the Shannon–Stratonovich formulation, the unknown system is $A$ with prior law $Q$, the decision is $B$, and the mutual-information constraint $I_S(A;B)\le \lambda$ limits the “bandwidth” or capacity of the observation channel from $A$ to $B$ [1405.5860]. In payoffs–beliefs duality, beliefs are points $p\in\Delta(K)$, actions are payoff vectors $a\in\mathbb R^K$, and the value function is the support function $V(p)=\sup_{a\in A}\langle a,p\rangle$ [1908.01633]. In hidden-variable models, the IV is the reduction in uncertainty about a current latent state produced by noisy past observations [2102.08841]. In feature selection, IV is explicitly identified with Jeffreys divergence, and the null hypothesis $H_0:\mathbf p=\mathbf q$ is equivalent to $\mathrm{IV}=0$ [2309.13183].

This suggests that “IV” is best understood not as a single invariant formula but as a task-indexed valuation operator: the object being valued can be experimental refinement, posterior dispersion, packet transmission, utterance novelty, or class separation.

## 2. Decision-theoretic value of information

The classical Shannon–Stratonovich theory starts from expected utility and adds an information resource constraint rather than abandoning the linearity of expected utility. With bounded utility $u(a,b)$ and joint law $P(A\cap B)=Q(A)\otimes P(B\mid A)$, the upper and lower branches are defined by
\[
\overline u_S(\lambda):=\sup_{P(B\mid A)}\{E_P[u(A,B)]:I_S(A;B)\le \lambda\},
\]
\[
\underline u_S(\lambda):=\inf_{P(B\mid A)}\{E_P[u(A,B)]:I_S(A;B)\le \lambda\},
\]
where
\[
I_S(A;B)=E_P\!\left[\ln\frac{dP(B\mid A)}{dP(B)}\right].
\]
The Stratonovich–Belavkin theorem states that $\overline u_S$ is nondecreasing, strictly increasing for $\lambda<\infty$, and concave in $\lambda$, whereas $\underline u_S$ is nonincreasing, strictly decreasing for $\lambda<\infty$, and convex in $\lambda$. If one defines
\[
U(\lambda)=
\begin{cases}
\overline u_S(\lambda), & \lambda\ge 0,\\[4pt]
\underline u_S(-\lambda), & \lambda\le 0,
\end{cases}
\]
then $U$ is S-shaped: concave for gains and convex for losses [1405.5860].

Within that framework, the asymmetry usually associated with prospect theory is derived without violating the independence axiom. The crucial claim is that if an expected-utility maximizer values information and therefore solves $\max E_P[u]$ subject to $I(P)\le \lambda$, then the S-shape follows from the linearity of expected utility plus the convexity of the information constraint. The paper states explicitly that no independence-axiom violation is needed [1405.5860]. A common misconception is therefore that gain/loss asymmetry necessarily requires a non-expected-utility model; in this formulation, it arises inside von Neumann–Morgenstern rationality once information itself is valued.

A complementary convex-analytic formulation is developed through payoffs–beliefs duality. For finite state space $K$, compact convex feasible payoff set $A\subset\mathbb R^K$, and belief $p\in\Delta(K)$, the best-reply payoff is
\[
V(p)=\sup_{a\in A}\langle a,p\rangle.
\]
An information structure is a $\Delta$-valued random posterior $q$ with $E[q]=p$, and the value of information is
\[
\mathrm{VoI}_A(q)=E[V(q)]-V(p)\ge 0.
\]
The subdifferential $\partial V(p)$ is the face of optimal actions at $p$, and the normal cone $N_A(a)$ is the set of beliefs for which $a$ is optimal. The confidence set
\[
C(p)=\bigcap_{a\in F_A(p)}N_A(a)\cap\Delta
\]
collects posteriors that preserve all actions optimal at the prior. The central theorem states
\[
\mathrm{VoI}_A(q)>0 \iff \mathbb P\{q\notin C(p)\}>0.
\]
Thus information has strictly positive value exactly when, with positive probability, it rules out at least one previously optimal action [1908.01633].

This duality also yields local and global estimates. There exist constants $\alpha,\beta>0$ such that
\[
\mathrm{VoI}_A(q)\le \alpha\,E[\mathrm{dist}(q,C(p))],\qquad
\mathrm{VoI}_A(q)\ge \beta\,\mathbb P\{q\notin U_\epsilon\},
\]
for $U_\epsilon=\{q\in\Delta:\mathrm{dist}(q,C(p))<\epsilon\}$ [1908.01633]. Near null information, the marginal value can be infinite, null, or positive and finite, depending on whether small posterior perturbations break a tie or remain in a smooth region of the value function.

A broader generalization replaces Shannon mutual information with a general leakage measure $\mathcal L(X\to Y)$ and loss $\ell$. The generalized VoI is
\[
\textsf V_{\mathcal L}^\ell(R;\mathcal Y):=
\sup_{p_{Y|X}:\mathcal L(X\to Y)\le R}\mathrm{gain}^\ell(X;Y),
\]
where
\[
\mathrm{gain}^\ell(X;Y)
=
\inf_{\delta^*}r(\delta^*,p_Y)-\inf_{\delta^*}r(\delta^*,p_{Y|X}).
\]
For standard losses, the upper bound follows from Bayes-risk definitions and the data-processing inequality, and in the classical loss case the “fundamental” VoI is achieved whenever the released variable is any sufficient statistic of the optimal synthetic output [2201.11449]. Because the paper interprets this optimization as a privacy–utility trade-off, IV here becomes a utility gain under a leakage budget rather than only a property of an experiment.

The Shannon/Stratonovich program has also been reformulated through entropy-constrained optimization over possibly infinite state and action spaces. With payoff $U(x,u)$ and KL budget $I$, one defines
\[
U(I):=\sup_{p(x,u)}\int U(x,u)\,p(x,u)\,dx\,du
\]
subject to
\[
D_{KL}(p(x,u)\|p(x)p(u))\le I,\qquad \int_U p(x,u)\,du=p(x).
\]
The Lagrangian first-order condition yields a Boltzmann–Gibbs form for the optimal joint law, and with the partition function $Z(\beta)$ and cumulant-generating function $T(\beta)=\ln Z(\beta)$, one obtains the parametric relations
\[
I(\beta)=\beta T'(\beta)-T(\beta),\qquad
U(\beta)=-T'(\beta),\qquad
V(I)=T(\beta)-\beta T'(\beta).
\]
The resulting single-number VoI is nonnegative, nondecreasing, concave, translation invariant under translation-invariant costs, and additive over independent subproblems [2303.16126].

## 3. Hidden states, status updates, and control-oriented IV

For latent-variable and hidden Markov models, IV is formulated as mutual information between the current source state and the observations available at the receiver. The general definition is
\[
v(t)=I(X_t;Y_{t_1'},Y_{t_2'},\dots,Y_{t_n'})
      =h(X_t)-h(X_t\mid Y_{t_1'},\dots,Y_{t_n'})\ge 0.
\]
Under the conditional independence assumption $X_t\perp\!\!\!\perp \mathbf Y\mid \mathbf X$, one has the bound
\[
v(t)\le \min\{I(X_t;\mathbf X),\,I(\mathbf X;\mathbf Y)\},
\]
and in the hidden Markov specialization the bound becomes a minimum of two sums of conditional mutual informations [2008.08907]. In the directly observed Markov case, the value reduces to $I(X_t;X_{t_n})$ [2102.08841].

For the Ornstein–Uhlenbeck process with additive Gaussian observation noise, the cited works derive closed-form expressions. In the single-observation case of the hidden Markov formulation,
\[
v(t)=I(X_t;Y_{t_n'})
     =-\frac12\ln\!\left(1-\frac{\gamma}{1+\gamma}e^{-2\kappa(t-t_n)}\right),
\]
where $\gamma=\frac{\mathrm{Var}(X_{t_i})}{\mathrm{Var}(N_{t_i'})}$ is an SNR-like ratio [2102.08841]. In the more general latent-variable treatment, the VoI is written through covariance determinants and the Matrix Determinant Lemma, with an explicit “noise-induced correction” to the pure OU Markov term [2008.08907]. Both treatments emphasize that AoI and VoI are not equivalent: AoI captures timeness, whereas VoI incorporates source correlation and observation noise. In the OU example, VoI decays exponentially in age, approximately as $O(e^{-2\kappa(t-t_n)})$, while AoI grows linearly [2008.08907].

In networked control, the term acquires a causal decision-theoretic meaning. One formulation introduces a packet-rate penalty
\[
R=\frac1{N+1}\mathbb E\!\left[\sum_{k=0}^N \ell(k)\sigma(k)\right],
\]
a regulation cost
\[
J=\frac1{N+1}\mathbb E\!\left[\sum_{k=0}^{N+1}x(k)^\top Q(k)x(k)+\sum_{k=0}^N u(k)^\top R(k)u(k)\right],
\]
and a team objective
\[
\Phi=\lambda R+J.
\]
The value of information at time $k$ is defined as the sensitivity of the encoder’s value function to forcing $\sigma(k)=0$ versus $\sigma(k)=1$:
\[
\mathrm{voi}(k,\mathcal I(k))
=
V(k,\mathcal I(k))|_{\sigma(k)=0}
-
V(k,\mathcal I(k))|_{\sigma(k)=1}.
\]
At equilibrium, it admits the closed form
\[
\mathrm{voi}(k,\mathcal I(k))
=
\tilde e(k)^\top A(k)^\top\Gamma(k+1)A(k)\tilde e(k)
-\theta(k)+\varrho(k),
\]
where $\tilde e(k)=\check x(k)-\hat x(k)$ is the estimation mismatch. The transmission rule is
\[
\sigma^\star(k)=\mathbf 1_{\mathrm{voi}(k,\mathcal I(k))\ge 0},
\]
and the equilibrium is stated to be globally optimal [2403.11927].

A closely related event-triggered control formulation defines
\[
\mathrm{voi}_k(\mathcal I^e_k)=V^e_k(\mathcal I^e_k)|_{\delta_k=0}-V^e_k(\mathcal I^e_k)|_{\delta_k=1},
\]
with controller law $u_k^\star=-L_k\hat x_k$ and Riccati recursion for $S_k$. At the Nash equilibrium,
\[
\mathrm{voi}_k(\mathcal I^e_k)
=
\tilde e_k^T A_k^T\Gamma_{k+1}A_k\tilde e_k-\theta_k+\varrho_k,
\]
and the optimal trigger is
\[
\delta_k^\star=\{\mathrm{voi}_k(\mathcal I^e_k)\ge 0\}.
\]
The paper emphasizes that $\mathrm{voi}_k$ is symmetric in the estimation mismatch and interprets it as benefit minus communication cost [1812.07534].

In remote estimation over an unreliable channel, IV is converted into a packet index. For packet $\psi$ at time $t$,
\[
W_\psi^2(t)=W_{s,\psi}^2(t)+W_{p,\psi}^2(t),
\]
with
\[
W_{s,\psi}^2(t)=a^2K^2a^{2\tau_\psi(t)}\sigma^2_{s,\psi},
\]
\[
W_{p,\psi}^2(t)=a^2K^2\sigma^2\frac{a^{2\tau_\psi(t)}-1}{a^2-1}+\sigma^2.
\]
The Value of Information is inversely ordered in $W^2$, equivalently
\[
\mathrm{VoI}_\psi(t):=-W_\psi^2(t),
\]
and the optimal policy is to transmit the packet with the largest current value of VoI, that is, the smallest $W^2$ [1908.01119]. The paper states that VoI decreases with age and increases with source precision, and concludes that a policy minimizing age of information does not necessarily maximize estimator performance.

## 4. Influence diagrams, Monte Carlo EVI, and nonmyopic approximation

In influence-diagram analysis, Information Value is usually expressed as expected value of perfect information (EVPI). For decision node $A$, chance node $X$, and utility $u(a,x)$, the expected utilities with and without perfect observation are
\[
EU(M)=\max_{a\in A}\sum_{i=1}^r p(x_i)\,u(a,x_i),
\]
\[
EU(M_{|X\ observed})=\sum_{i=1}^r p(x_i)\max_{a\in A}u(a,x_i),
\]
and
\[
EVPI_M(A\mid X)=EU(M_{|X\ observed})-EU(M).
\]
The net quantity is
\[
NEVPI_M(A\mid X)=EVPI_M(A\mid X)-Cost(X).
\]
Graph-theoretic analysis then yields qualitative dominance relations. If $(X\perp V\mid A)$ in the influence diagram, then $EVPI_M(A\mid X)=0$. If $Y\perp V\mid X$ and neither $X$ nor $Y$ is a descendant of $A$, then
\[
EVPI_M(A\mid X)\ge EVPI_M(A\mid Y).
\]
These results permit a nonnumerical partial order of variables by informational relevance using canonical form and d-separation alone [1302.3596].

Monte Carlo decision models motivate an approximate preposterior computation of EVI. Let $d^*$ be the Bayes-optimal prior decision and $d_e^*$ the posterior-optimal decision after evidence $e$. Then
\[
\mathrm{EVI}(e)
=
E[v(X,d_e^*)\mid S]-E[v(X,d^*)\mid S].
\]
With the regret variable
\[
Z=v(X,d^*)-v(X,d^+),
\]
where $d^+$ is the second-best prior decision, the perfect-information value is
\[
\mathrm{EVPI}=\int_{-\infty}^0 |z|\,f_Z(z)\,dz.
\]
The paper introduces a linear approximation
\[
v(X,d_i)\approx \alpha_i+\sum_{j=1}^n \beta_{ij}X_j,
\]
uses multiple linear regression to estimate the coefficients from Monte Carlo samples, and derives Gaussian preposterior formulas for perfect or partial information on individual variables or subsets [1302.6794]. The stated computational advantage is that, once the surrogate is fit, EVI for different information sets reduces to variance adjustments and a one-dimensional normal integral.

Approximate nonmyopic computation addresses the intractability of evaluating all possible sequences of observations. In a binary-hypothesis diagnosis setting with tests $E_1,\dots,E_n$, the myopic policy computes
\[
\mathrm{NVOI}(E)=\mathrm{VOI}(E)-C(E)
\]
under the assumption that at most one additional test will be performed. The nonmyopic approximation instead orders tests by myopic NVOI, considers prefixes $S_m=\{E_1,\dots,E_m\}$, and approximates the total log-odds weight
\[
W=\sum_{i=1}^m w_i,\qquad
w_i=\ln\frac{P(E_i\mid H)}{P(E_i\mid \neg H)},
\]
by a normal distribution through the central-limit theorem. The threshold
\[
W^*=\ln\frac{p^*}{1-p^*}-\ln O(H)
\]
determines the action, and the resulting approximation yields $\mathrm{NVOI}(S_m)$ for each prefix [1303.5720]. The methodological significance is that nonmyopic value can be approximated in polynomial time rather than by exact enumeration of exponentially many observation sequences.

## 5. Utterance predictability and psychometric IV

In dialogue and text comprehension, Information Value is defined at the utterance level rather than the token level. Let $X$ be the preceding discourse context, $Y$ the next utterance, and
\[
A_x=\{a_1,a_2,\dots,a_N\}
\]
a hypothetical set of plausible next utterances after context $x$. The Information Value of observing $Y=y$ in context $X=x$ is
\[
I(Y=y\mid X=x)=d(y,A_x),
\]
where $d$ is a distance metric between the target utterance and the alternative set. Two scalar summaries are used:
\[
I_{\text{mean}}(y\mid x)=\frac1{|A_x|}\sum_{a\in A_x} d(y,a),
\qquad
I_{\min}(y\mid x)=\min_{a\in A_x} d(y,a).
\]
Smaller values mean that the utterance is closer to expectations and therefore more predictable [2310.13676].

Because a human comprehender’s full alternative set is not directly observable, plausible alternatives are proxied by samples from neural language models conditioned on the context. The paper reports GPT-2, DialoGPT, GPT-Neo, and OPT model families; dialogue models are fine-tuned on Switchboard or DailyDialog, whereas text models are used off-the-shelf. It studies 11 decoding strategies: ancestral sampling, temperature sampling with $\alpha\in\{0.75,1.00,1.25\}$, nucleus sampling with $p\in\{0.8,0.85,0.9,0.95\}$, and locally-typical sampling with $\tau\in\{0.2,0.3,0.85,0.95\}$. Distances include lexical $n$-gram overlap, syntactic POS-tag $n$-gram distance, and semantic distance via cosine or Euclidean distance of SBERT sentence embeddings [2310.13676].

The central contrast is with utterance-level aggregations of token surprisal, such as mean, total, max, variance, or superlinear weighting of
\[
I(u)=-\log_2 p(u\mid \text{context}).
\]
The paper argues that aggregated surprisal conflates lexical, syntactic, and semantic unpredictability into one bit-count and overestimates the surprise of paraphrases that compete for probability mass under softmax, whereas IV measures how far a full utterance is from plausible alternatives in interpretable dimensions [2310.13676].

Empirically, IV is reported as a stronger predictor of contextual acceptability than token-based surprisal in dialogue, and as complementary to surprisal for reading times. On Switchboard acceptability, the best IV variant is semantic-min with Spearman $\rho=-0.702$, compared with utterance-max surprisal at $\rho=-0.506$. On DailyDialog acceptability, semantic-min IV attains $\rho=-0.584$, compared with superlinear surprisal ($k=2.5$) at $\rho=-0.457$. On Provo reading times, syntactic-min IV yields $\rho=+0.421$, while surprisal variance gives $\rho=+0.495$. Joint models also improve fit: in Switchboard, adding semantic IV to surprisal raises log-likelihood from $6.63$ to $34.37$, and in Provo, adding syntactic IV increases log-likelihood by $+16.66$ from $59.04$ to $75.70$ [2310.13676].

A frequent misunderstanding would be to treat this IV as a replacement for surprisal. The paper states instead that it is complementary to surprisal for predicting eye-tracked reading times and that semantic IV dominates in dialogue while lexical or syntactic IV matter more in reading [2310.13676]. The construct is therefore not a reparameterized surprisal score, but a set-distance measure over utterance alternatives.

## 6. IV as Jeffreys divergence in feature selection

In binary supervised learning, Information Value is a divergence-based measure of class separation after discretizing a predictor. Let $Y\in\{0,1\}$ and let the support of predictor $X$ be partitioned into bins $a_1,\dots,a_r$. Define
\[
p_j=\Pr\{X=a_j\mid Y=1\},\qquad
q_j=\Pr\{X=a_j\mid Y=0\}.
\]
Then
\[
\mathrm{IV}
=
J(\mathbf p,\mathbf q)
=
\sum_{j=1}^r (p_j-q_j)\ln\frac{p_j}{q_j},
\]
which the paper identifies as the Jeffreys divergence, that is, the symmetric Kullback–Leibler divergence between the two class-conditional distributions [2309.13183]. The empirical version replaces $p_j$ and $q_j$ by sample proportions $\hat p_j$ and $\hat q_j$ computed from the counts of positive and negative labels in each bin.

The same work develops a non-parametric hypothesis test for predictive power. The null and alternative hypotheses are
\[
H_0:\mathbf p=\mathbf q\quad(\mathrm{IV}=0),
\qquad
H_A:\mathbf p\neq \mathbf q\quad(\mathrm{IV}>0).
\]
The test statistic is the sample IV,
\[
T=\widehat{\mathrm{IV}}
=
\sum_{j=1}^r (\hat p_j-\hat q_j)\ln\frac{\hat p_j}{\hat q_j},
\]
and, under mild regularity conditions, the normalized statistic
\[
Z
=
\sqrt{\frac{nm}{(n+m)\widehat\Sigma_{n,m}}}\,
\widehat{\mathrm{IV}}
\]
converges in distribution to $\mathcal N(0,1)$ under $H_0$, with $Z^2\approx \chi_1^2$ [2309.13183]. This yields p-values and significance thresholds that depend on sample size, class imbalance, and plug-in variance estimates.

The paper is explicitly critical of fixed threshold heuristics such as $\mathrm{IV}>0.1$. It notes that these practical criteria are “mysterious and lacking theoretical arguments,” ignore sample size, do not adapt to class imbalance, and provide no control of Type I error or power [2309.13183]. The proposed J-Divergence test is reported to be more reliable, particularly in unbalanced data sets.

Simulation and case-study results reinforce that distinction. Under mild imbalance with $n=300$, $m=50\,000$, $r=10$, and $\alpha=0.1\%$, the J-Divergence test maintains Type I approximately $0.001$ and reaches power approaching $1$ as the distributions diverge, whereas the rule $\mathrm{IV}>0.1$ has essentially zero power until divergence is large. In varying imbalance scenarios, the test’s power curves remain stable for $\tfrac n{n+m}\in[0.01,0.5]$, while the threshold rule can have Type I up to $40\%$ when $\tfrac n{n+m}$ is small [2309.13183].

In the Vesta e-commerce fraud data with 369 features, LightGBM trained after J-Divergence selection used 262 features and achieved precision $0.88491$, recall $0.73812$, AUC $0.97188$, and F1 $0.79734$. The threshold rule $\mathrm{IV}>0.1$ selected 220 features and achieved precision $0.87141$, recall $0.73457$, AUC $0.97112$, and F1 $0.79666$, while using no selection retained all 369 features with precision $0.85496$, recall $0.74594$, AUC $0.97098$, and F1 $0.79624$ [2309.13183]. The library `statistical-iv` accompanies that methodology.

Across these literatures, IV is unified less by a single formula than by a common role: it quantifies how much a signal, observation, experiment, utterance, or predictor matters for a downstream inferential or decision problem. The precise quantity depends on the underlying model class. In expected-utility theory it is a constrained gain frontier; in convex decision geometry it is the support-function gain from posterior variation; in latent-variable systems it is uncertainty reduction; in control it is benefit minus communication cost; in NLP it is distance from plausible alternatives; and in feature selection it is a symmetric divergence between class-conditional distributions.

Source: https://www.emergentmind.com/topics/information-value-iv