Papers
Topics
Authors
Recent
Search
2000 character limit reached

Principle of Available Information (PAI)

Updated 10 July 2026
  • PAI is a principle asserting that conditioning on additional, probabilistically modeled evidence never increases the expected Bayes risk.
  • The formal framework leverages Bayesian updating, the concavity of the Bayes risk function, and likelihood-ratio manipulations to demonstrate decision improvement in expectation.
  • PAI interprets differences in likelihood-ratio outcomes as results of varying available information sets, thereby unifying disparate inferential practices.

The Principle of Available Information (PAI) is the Bayesian-decision-theoretic claim that additional information, when incorporated by appropriate probabilistic updating, cannot worsen optimal decisions on average. In its clearest recent formalization, a prior hypothesis distribution π\pi is updated to a posterior π(E)\pi(E) after observing information EE, and the Bayes risk satisfies

E[c(π(E))]c(π).\mathbb{E}[c(\pi(E))]\le c(\pi).

Within this formulation, “available information” means whatever observable information can be conditioned on and modeled—raw observations, feature vectors, reduced scores, or augmented conditioning variables. PAI therefore functions both as a theorem about Bayes risk and as an interpretive principle: apparently different inferential procedures are often best understood as the same inferential logic applied to different information sets (Boonstra et al., 2024).

1. Formal statement and decision-theoretic content

PAI is naturally expressed in a finite Bayesian decision problem. Let H1,,HnH_1,\dots,H_n be mutually exclusive and exhaustive hypotheses with prior probability vector

π=(π1,,πn):=(P(H1),,P(Hn)),\pi=(\pi_1,\dots,\pi_n):=\big(P(H_1),\dots,P(H_n)\big),

let A1,,AmA_1,\dots,A_m be possible actions, and let cijc_{ij} be the cost of choosing action AiA_i when HjH_j is true. For any probability vector π(E)\pi(E)0, the expected cost of action π(E)\pi(E)1 is

π(E)\pi(E)2

A Bayes decision π(E)\pi(E)3 minimizes this quantity, and the corresponding minimal expected cost is denoted

π(E)\pi(E)4

If π(E)\pi(E)5 is new evidence with realization π(E)\pi(E)6, posterior updating yields

π(E)\pi(E)7

PAI is then the inequality

π(E)\pi(E)8

which states that the average Bayes cost after conditioning on available information is at most the Bayes cost before conditioning (Boonstra et al., 2024).

The principle is weakly monotone rather than strictly monotone. It guarantees “decrease or stay the same,” not universal strict improvement. Equality arises in “boundary cases where no or not enough new information is incorporated,” which includes cases where the added variable is redundant for the decision problem or too weak to alter the Bayes-optimal action. The operative notion of improvement is therefore expected decision quality, not pointwise dominance in every realized case.

This formulation sharply separates two issues that are often conflated. First, there is the normative question of what one should do if information is available and probabilistically modelable; PAI answers that such information should be incorporated. Second, there is the representational question of what counts as information; here the framework is deliberately broad. The evidence variable may be a full observation π(E)\pi(E)9, a reduced representation EE0, or an augmented tuple such as EE1. PAI does not privilege raw data as such; it privileges coherent conditioning on the most informative available representation that is actually used.

2. Likelihood-ratio mechanism and proof structure

In the two-hypothesis case, PAI can be expressed directly in likelihood-ratio form. With hypotheses EE2, prior odds

EE3

conditioning information EE4, and evidence EE5, the likelihood ratio is

EE6

so posterior odds satisfy

EE7

A binary decision problem with asymmetric costs makes the mechanism especially transparent. If costs are normalized so that

EE8

then the Bayes decision chooses EE9 iff

E[c(π(E))]c(π).\mathbb{E}[c(\pi(E))]\le c(\pi).0

After observing E[c(π(E))]c(π).\mathbb{E}[c(\pi(E))]\le c(\pi).1, the decision switches when the posterior odds cross the threshold, equivalently when

E[c(π(E))]c(π).\mathbb{E}[c(\pi(E))]\le c(\pi).2

The proof controls the probability of such threshold crossings under each hypothesis by elementary likelihood-ratio identities, yielding the conclusion that the expected posterior Bayes cost cannot exceed the prior Bayes cost (Boonstra et al., 2024).

The fully general proof does not rely on the binary case. Its central ingredient is the concavity of the Bayes-risk function E[c(π(E))]c(π).\mathbb{E}[c(\pi(E))]\le c(\pi).3 on the probability simplex: E[c(π(E))]c(π).\mathbb{E}[c(\pi(E))]\le c(\pi).4 Taking E[c(π(E))]c(π).\mathbb{E}[c(\pi(E))]\le c(\pi).5 and E[c(π(E))]c(π).\mathbb{E}[c(\pi(E))]\le c(\pi).6 gives

E[c(π(E))]c(π).\mathbb{E}[c(\pi(E))]\le c(\pi).7

The remainder reduces to the identity

E[c(π(E))]c(π).\mathbb{E}[c(\pi(E))]\le c(\pi).8

which is the law of total expectation for posterior probabilities. The distinctive feature of the recent proof is that it establishes this through likelihood-ratio manipulations rather than through auxiliary proper-scoring-rule machinery (Boonstra et al., 2024).

The same work also clarifies the relation to proper scoring rules rather than rejecting it. Any Bayes decision problem induces a proper scoring rule via

E[c(π(E))]c(π).\mathbb{E}[c(\pi(E))]\le c(\pi).9

where H1,,HnH_1,\dots,H_n0 is an action minimizing expected cost under H1,,HnH_1,\dots,H_n1. In this sense, Bayes decision cost and proper scoring are two views of the same object. PAI can therefore be read either as a theorem about Bayes risk or as a theorem about proper scoring under conditioning, but its recent formulation makes the likelihood-ratio structure explicit.

3. Information sets rather than competing “LR systems”

A central interpretive consequence of PAI is that many disputes in forensic likelihood-ratio practice are better understood as disputes about information sets rather than about inferential logics. The key claim is that “there is no such thing as different ‘LR-systems’, there are only differences in the available information.” On this view, feature-based versus score-based LRs, and common-source versus specific-source LRs, address the same substantive hypotheses while conditioning on different subsets of the evidential situation (Boonstra et al., 2024).

This perspective follows directly from the posterior-odds identity. Changing the LR usually means changing either the conditioning information H1,,HnH_1,\dots,H_n2 or the evidence variable being processed. If one computes an LR from a score H1,,HnH_1,\dots,H_n3 rather than from the full observation H1,,HnH_1,\dots,H_n4, the inferential structure is not altered; only the sigma-field of conditioning is reduced. Likewise, if one adds source-specific parameters H1,,HnH_1,\dots,H_n5 to the conditioning set, the LR changes because the information position has changed, not because a different epistemic framework has been adopted.

PAI therefore treats the choice of LR as a question of evidential representation. If only H1,,HnH_1,\dots,H_n6 is available, the appropriate LR is the LR for H1,,HnH_1,\dots,H_n7. If the full H1,,HnH_1,\dots,H_n8 is available and modelable, the LR for H1,,HnH_1,\dots,H_n9 is weakly preferable in expected decision cost. If further source parameters are available, conditioning on them yields another weak improvement. The principle is not that one must always use raw data, but that one should not confuse reduced-information analysis with a distinct inferential system.

This interpretation also resolves a recurrent misconception. Score-based or otherwise reduced-information analyses are not rendered incoherent merely because richer information exists in principle. They remain valid for the reduced information set they actually represent. What PAI adds is a ranking: relative to fuller admissible information, such analyses are suboptimal from the standpoint of expected Bayes cost.

4. Feature reduction, score-based analysis, and source conditioning

The score-based versus feature-based issue is the most explicit operational instance of PAI. Let π=(π1,,πn):=(P(H1),,P(Hn)),\pi=(\pi_1,\dots,\pi_n):=\big(P(H_1),\dots,P(H_n)\big),0 be the full evidence and let π=(π1,,πn):=(P(H1),,P(Hn)),\pi=(\pi_1,\dots,\pi_n):=\big(P(H_1),\dots,P(H_n)\big),1 be a reduced score or feature map. The reduced variable π=(π1,,πn):=(P(H1),,P(Hn)),\pi=(\pi_1,\dots,\pi_n):=\big(P(H_1),\dots,P(H_n)\big),2 is genuinely informative, even though it is less informative than π=(π1,,πn):=(P(H1),,P(Hn)),\pi=(\pi_1,\dots,\pi_n):=\big(P(H_1),\dots,P(H_n)\big),3. PAI yields both

π=(π1,,πn):=(P(H1),,P(Hn)),\pi=(\pi_1,\dots,\pi_n):=\big(P(H_1),\dots,P(H_n)\big),4

and, by iterating the same logic,

π=(π1,,πn):=(P(H1),,P(Hn)),\pi=(\pi_1,\dots,\pi_n):=\big(P(H_1),\dots,P(H_n)\big),5

Accordingly, decisions based on scores are an improvement over ignoring the evidence entirely, while decisions based on the full evidence are, on average, an improvement over decisions based only on the score (Boonstra et al., 2024).

The same structure governs common-source versus specific-source analyses. A common-source LR typically compares

π=(π1,,πn):=(P(H1),,P(Hn)),\pi=(\pi_1,\dots,\pi_n):=\big(P(H_1),\dots,P(H_n)\big),6

using observations such as π=(π1,,πn):=(P(H1),,P(Hn)),\pi=(\pi_1,\dots,\pi_n):=\big(P(H_1),\dots,P(H_n)\big),7 when source parameters are not fully known. A specific-source LR arises when one additionally knows, for example, π=(π1,,πn):=(P(H1),,P(Hn)),\pi=(\pi_1,\dots,\pi_n):=\big(P(H_1),\dots,P(H_n)\big),8, so that the relevant evidence becomes π=(π1,,πn):=(P(H1),,P(Hn)),\pi=(\pi_1,\dots,\pi_n):=\big(P(H_1),\dots,P(H_n)\big),9 or equivalently A1,,AmA_1,\dots,A_m0 with redundancy removed. PAI interprets this not as a switch of model class or inferential doctrine, but as the same hypothesis test under a richer information set.

A concrete illustration is the DNA kinship example in which

A1,,AmA_1,\dots,A_m1

If only the profiles A1,,AmA_1,\dots,A_m2 are known, the LR is

A1,,AmA_1,\dots,A_m3

If the parents’ profiles A1,,AmA_1,\dots,A_m4 later become known, the LR becomes

A1,,AmA_1,\dots,A_m5

In the latter representation, A1,,AmA_1,\dots,A_m6 becomes redundant because the relevant uncertainty is now whether A1,,AmA_1,\dots,A_m7 could be a child of the known parents. The same example admits a score-based reduction A1,,AmA_1,\dots,A_m8, such as the number of shared alleles. The simulated comparisons reported for these information positions follow the PAI ordering: using the score improves over using no data, using the full profiles improves over using only the score, and adding the parents’ profiles improves further (Boonstra et al., 2024).

The principle is accompanied by probabilistic control rather than absolute guarantees for individual cases. For two hypotheses, the bounds

A1,,AmA_1,\dots,A_m9

show that strong reversals in the wrong direction are controlled in probability when additional conditioning information is added. This does not preclude reversal in a single realization; it states that average and tail behavior favor richer coherent conditioning.

PAI has both normative and positive neighbors. A closely related but distinct positive theory asks not whether more available information should be used, but what information structure could have generated observed behavior. In “Revealed Information,” the analyst observes only an unconditional action distribution cijc_{ij}0, not state-contingent choices, and asks when it can be rationalized as if a decision maker first received information and then chose optimally. The pair cijc_{ij}1 is BCE-consistent given utility cijc_{ij}2 iff there exists a joint distribution satisfying obedience and marginal conditions, equivalently iff the prior lies in the weighted Minkowski sum

cijc_{ij}3

where cijc_{ij}4 is the set of posterior beliefs under which action cijc_{ij}5 is optimal (Doval et al., 2024).

This revealed-information framework is not PAI in the normative sense, but it sharpens the epistemic background against which PAI operates. It shows that observed behavior generally does not identify the exact information available to the decision maker; rather, it identifies a compatibility set of priors, posterior distributions, or information structures. The support-function characterization,

cijc_{ij}6

gives a testable criterion for whether some latent information structure can rationalize the action frequencies. A plausible implication is that PAI’s normative prescription should be distinguished from empirical claims about what information an agent in fact possessed.

A second adjacent line of work appears in evidence-combination theory. “Combination of Evidence Using the Principle of Minimum Information Gain” proposes that all available constraints—marginals, compatibility exclusions, and any known conditionals—should be imposed, and that among the admissible joint distributions one should choose the one with minimum information gain, equivalently maximum entropy or minimum cross-entropy relative to the product-of-marginals baseline (1304.1135). The information-gain term is

cijc_{ij}7

which the paper rewrites as

cijc_{ij}8

This is not the same doctrine as the Bayes-risk PAI theorem, but it is conceptually allied: both frameworks insist that inference should encode exactly the information actually available and avoid unsupported supplementation.

6. Scope, caveats, and terminological disambiguation

PAI is sometimes misunderstood as a blanket claim that more information always improves every realized decision. The formal result is weaker and more precise. Improvement is in expectation. In a particular case, conditioning on additional information can lead to a posterior Bayes decision with higher realized cost than the action that would have been chosen without that information. The theorem states only that the average over evidence realizations is non-increasing. It also presumes “appropriate update of the probabilities of the hypotheses,” that is, coherent Bayesian conditioning under an adequate probabilistic model (Boonstra et al., 2024).

This modeling caveat is substantive. PAI does not license arbitrary LR computations detached from the actual information structure. Nor does it require using every measurable variable merely because it exists. The relevant condition is that the added information be available in the evidential situation and justifiably modelable. The recent forensic treatment explicitly notes that a more accurate model for simplified data may be preferable to a flawed model for fuller data. PAI supports using the best-justified available information, not mechanically maximizing dimensionality.

A further practical caveat concerns reporting. One should not compute multiple LRs from different information sets and then report the most discriminatory one. The LR should correspond to the actual knowledge position being represented. This restriction is intrinsic to the principle: the inferential object is indexed by the information set, so post hoc selection across incomparable information positions distorts the evidential statement.

Finally, the acronym “PAI” is overloaded in the contemporary arXiv literature. In unrelated work it denotes Alibaba Cloud’s “Platform of AI” (Hu et al., 2017), “Preserving Amplitude Information” in time-series anomaly detection (Zhang et al., 8 Jun 2026), and “Probabilistic Angle Interpolation” in quantum simulation (Kiumi et al., 2024). These usages are terminological collisions rather than substantive connections. Within the present topic, PAI denotes the principle that coherent exploitation of available information weakly improves optimal decision-making on average, and that many methodological disputes are most parsimoniously understood as disputes over information sets rather than over inferential principle.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Principle of Available Information (PAI).