---
title: Minimum Updating Principle
url: https://www.emergentmind.com/topics/minimum-updating-principle
type: topic
---

# Minimum Updating Principle

Searching arXiv for recent and foundational papers related to the Minimum Updating Principle.
The minimum updating principle is a family of doctrines in epistemology, probability, belief revision, and database theory according to which an existing informational state should be revised only insofar as newly received information requires, and no further. Across the literature, that common intuition is realized in materially different ways: by maximizing relative entropy subject to posterior constraints, by preserving conditional structure under Bayesian conditioning, by minimizing information gain in evidence combination, by selecting inclusion-minimal base changes in knowledge bases, or by restricting update rules through calibration, rectangularity, or ambiguity-sensitive axioms. A recurrent conclusion is that there is no single universally correct update rule; what counts as “minimal” depends on the representation of uncertainty, the semantics of the evidence, and the structural assumptions imposed on the observation or revision process [0708.1593] [1303.5752] [1407.7183].

## 1. Conceptual meaning and scope

In the entropy-based literature, the principle is stated directly: one should change prior beliefs only where corrective new evidence has been supplied. Giffin and Caticha formulate this by maximizing relative entropy on the joint space of data and parameters, so that the posterior is the distribution closest to the prior while satisfying the new constraints [0708.1593]. In belief-function theory, Smets treats updating as a family of conditioning schemes, each preserving a different aspect of the prior state, and explicitly argues that there is no single “good” rule because different rules correspond to different interpretations of what the evidence means [1303.5752]. In the study of naive and sophisticated probability spaces, the same intuition is sharply qualified: minimal-change rules such as conditioning, Jeffrey conditioning, or minimum relative entropy are justified only when the observation mechanism has the right structure; otherwise they may be conservative relative to the wrong representation of evidence [1407.7183].

The principle is therefore best understood as a meta-principle rather than a single algorithm. In one strand, “minimum” means minimizing a divergence functional. In another, it means preserving relative probabilities or preserving support only within subsets. In database theory, it means changing only the updatable extensional facts while leaving the rule layer and integrity constraints fixed. In ambiguity models, it can mean avoiding dilation, imposing maximal dynamic consistency, or staying as close as possible to a benchmark posterior subject to an evidence-selection rule. This suggests that the principle is representation-relative: the relevant invariants are different for probabilities, belief functions, credal sets, Horn knowledge bases, and ambiguity-sensitive preference models.

A persistent misconception is that minimal updating always coincides with ordinary conditioning. The cited literature does not support that claim. Conditioning is one special case of minimal revision, but only when the semantics of the evidence and the structure of the state space warrant it. In other settings, the appropriate minimal revision may be Jeffrey conditioning, maximum-entropy updating, specialization, imaging, minimal hitting-set change, or even ignoring the observation.

## 2. Bayesian and entropy-based formulations

The most explicit information-theoretic realization appears in the method of Maximum relative Entropy. Giffin and Caticha maximize
$$
S[P,P_{\text{old}}] = -\int dx\,d\theta\; P(x,\theta)\log\frac{P(x,\theta)}{P_{\text{old}}(x,\theta)}
$$
over posterior joint distributions subject to normalization and whatever posterior constraints express the new information. When the only new information is observed data \(x'\), represented as
$$
P(x)=\int d\theta\,P(x,\theta)=\delta(x-x'),
$$
the method yields Bayes’ rule exactly:
$$
P_{\text{new}}(\theta)=P_{\text{old}}(\theta\mid x').
$$
When data and a moment constraint
$$
\int dx\,d\theta\,P(x,\theta)f(\theta)=F
$$
are imposed simultaneously, the posterior takes the canonical form
$$
P_{\text{new}}(\theta)=P_{\text{old}}(\theta\mid x')\frac{e^{\beta f(\theta)}}{Z},
$$
with \(\beta\) determined by \(\frac{d\log Z}{d\beta}=F\). This framework interprets minimal updating as maximum preservation of the prior joint structure consistent with the active constraints, and it distinguishes sequential from simultaneous processing of constraints according to whether earlier constraints have been superseded or remain valid [0708.1593].

Wong and Lingras formulate a closely related doctrine as the “principle of minimum information gain” for combining evidence. Given marginal evidence on two frames, they seek a joint distribution \(P(\{(s,s')\})\) that preserves the marginals and compatibility constraints while minimizing
$$
\sum_{(s,s')\in S\times S'} P(\{(s,s')\}) \log \frac{P(\{(s,s')\})}{P(\{s\})P(\{s'\})}.
$$
They show that this criterion is equivalent to maximum entropy and to a special case of minimum cross-entropy. It reduces to Bayes’ rule when all conditional probabilities are available, and to Dempster’s rule when the normalization constant is equal to one. Their interpretation of “minimum” is precise: preserve the product-of-marginals baseline unless the constraints force dependence, and preserve the original marginals rather than distorting them by normalization [1304.1135].

A more recent KL-based result concerns Jeffrey’s rule. The paper “Jeffrey’s update rule as a minimizer of Kullback-Leibler divergence” does not prove the classical constrained \(I\)-projection theorem. Its main theorem is instead that if
$$
\theta_{t+1}:=\overleftarrow C_{\theta_t}(\tau),
$$
then
$$
D_{\mathrm{KL}(\tau \,\|\, \overrightarrow C(\theta_{t+1})) \le D_{\mathrm{KL}(\tau \,\|\, \overrightarrow C(\theta_t)).
$$
Within its EM-style proof, Jeffrey’s posterior is also shown to be the unique minimizer of
$$
D_{\mathrm{KL}\!\big(\overleftarrow C_{\theta_t}(\tau)\,\|\,\theta\big)}
$$
over the simplex of latent-state distributions. Thus the paper supplies a minimum-change characterization, but in observation space and in an auxiliary reverse-KL problem, not as a general theorem that Jeffrey conditioning minimizes posterior-to-prior KL subject to marginal constraints [2502.15504].

These formulations share a common structure. The prior is not discarded; it becomes the reference measure or baseline geometry. New information is encoded as explicit constraints, and the posterior is the admissible state that departs least from that baseline according to the relevant divergence. At the same time, the formal object being preserved differs across frameworks: a joint prior on \((x,\theta)\), a product-of-marginals distribution on an evidence frame, or the prediction induced by a latent prior through a channel.

## 3. Conditioning rules, belief functions, and naive-versus-sophisticated evidence

In belief-function theory, the minimum updating intuition fragments into several distinct conditioning schemes. Smets surveys unnormalized Dempster conditioning, normalized Dempster conditioning, Bayesian conditioning as a special case, Yager–Kohlas conditioning, the geometric rule, specialization, imaging, and upper/lower Bayesian conditioning. These rules differ in what they preserve when evidence eliminates a set \(\bar A\): unnormalized conditioning transfers mass by \(X\mapsto X\cap A\) and leaves conflict on \(\varnothing\); normalized conditioning rescales surviving masses under a closed-world assumption; Yager–Kohlas reallocates conflict to \(A\); the geometric rule preserves only mass already attached to subsets of \(A\); specialization redistributes each mass \(m(X)\) only to subsets \(B\subseteq X\); and imaging redistributes support to “closest” surviving states through a transition kernel. Smets’ central claim is that the correct rule depends on what the evidence means, not merely on the syntax of conditioning [1303.5752].

The distinction between naive and sophisticated spaces sharpens this point in probabilistic form. Halpern and collaborators distinguish a naive space \(W\) of worlds from a sophisticated space \(\mathcal R\) that also encodes how the observation arose. For ordinary conditioning, naive-space updating by \(\Pr_W(\cdot\mid U)\) matches sophisticated conditioning only under CAR, the coarsening-at-random condition. For Jeffrey conditioning, the analogous correctness criterion is a generalized CAR condition. For MRE, minimum cross-entropy, or minimum information gain, the paper proves a negative result: except in special cases reducible to Jeffrey conditioning, there is no comparable general structural condition guaranteeing that naive-space MRE matches conditioning in the sophisticated space. The paper explicitly concludes that minimal information gain in the naive space is not by itself a justification for correct updating [1407.7183].

Updating with sets of probabilities yields a further refinement. Grünwald and Halpern analyze finite decision problems with uncertainty represented by a set \(\mathcal P\subseteq \Delta(\mathcal X\times\mathcal Y)\) under the minimax criterion. They define general update rules \(\Pi(\mathcal P,x)\), including ordinary conditioning, ignoring information, and \(\mathcal C\)-conditioning based on a partition \(\mathcal C\) of \(\mathcal X\). In the \(\mathcal P\)-\(X\)-game, where the bookie chooses after observing \(x\), conditioning on \(X=x\) is a posteriori minimax optimal. In the \(\mathcal P\)-game, where the bookie chooses before \(x\), conditioning can fail, and even ignoring the observation can be a priori minimax optimal under a stated independence condition. The key structural condition is rectangularity, written as \(\mathcal P=\langle\mathcal P\rangle\). Under that condition, ordinary conditioning becomes weakly time consistent, and with conservativeness it becomes time consistent; under convexity and rectangularity it is also sharply calibrated [1401.3906].

Taken together, these results show that “minimal” conditioning is meaningful only relative to the correct evidence model. In belief functions, different transfer rules preserve different structural commitments. In imprecise-probability models, a more conservative update can be required by minimax considerations. In naive-space probability theory, even highly conservative entropy minimization can be wrong if the observation protocol has been omitted from the state description.

## 4. Minimal change in knowledge bases and relational view updating

In database theory, the minimum updating principle is formalized as minimal disturbance of the updatable base under immutable rules and integrity constraints. Delhibabu and Lakemeyer represent a Horn knowledge base as
$$
KB = KB_I \cup KB_U \cup KB_{IC},
$$
with \(KB_I\) immutable, \(KB_U\) updatable, and \(KB_{IC}\) the integrity constraints. For deductive databases this becomes
$$
DDB = \langle IDB, EDB, IC\rangle.
$$
A view update is a request to change the truth of a derived atom by modifying only relevant base facts. Minimal realizations are those in which no constituent base update can be removed while still realizing the request. The paper makes minimality precise in several connected ways: inclusion-minimal abductive explanations, kernel-style change, minimal hitting sets, and minimal models of transformed disjunctive programs. Cardinality minimality is explicitly not the main criterion; the dominant criterion is set inclusion [1407.3512].

The paper adapts AGM-style belief revision to what it calls knowledge-base dynamics or base dynamics. Revision \(KB*\alpha\) is constrained by postulates including Closure, Weak Success, Inclusion, Immutable-inclusion, Vacuity 1, Vacuity 2, Consistency, Preservation, and three relevance postulates, of which relevance and weak relevance are the ones most directly tied to minimal updating. The key idea is that only those base facts implicated in the conflict with \(\alpha\) may be altered. Formally, the revision process identifies minimal supports or offending subsets and computes inclusion-minimal hitting sets:
$$
HS \subseteq \bigcup S,\qquad R\cap HS\neq\emptyset \text{ for every non-empty }R\in S.
$$
Changing a hitting set guarantees success because every derivation or conflict path is intercepted; choosing a minimal hitting set guarantees that no unnecessary base fact is altered [1407.3512].

This abstract account is connected to concrete algorithms. The paper transforms the database and update request into a disjunctive datalog program and uses hyper tableaux for deletions and magic sets for insertions. For deletion, an open finished branch \(b\) of the update tableau yields
$$
HS(b)=\{A\in EDB\mid \neg A\in b\},
$$
and strong minimality is tested by
$$
\forall s\in HS(b):\; IDB \cup EDB \backslash HS(b) \cup \{s\} \vdash A.
$$
For insertion, a symmetric condition is imposed on close finished branches in the magic-set construction. The paper explicitly proves that the strong minimality test and the groundedness test are equivalent. It then proves that Algorithms 3 and 4 satisfy \((KB*1)\)–\((KB*6)\) and the strong relevance postulate \((KB*7.1)\), while the materialized-view Algorithms 5 and 6 satisfy \((KB*1)\)–\((KB*6)\) and weak relevance \((KB*7.3)\) [1407.3512].

Here the minimum updating principle is not metaphorical. It is a formally axiomatized requirement that only the necessary \(EDB\) facts be changed, under a fixed \(IDB\) and fixed \(IC\), and it is realized algorithmically through inclusion-minimal hitting-set computation over abductive explanations and kernels.

## 5. Ambiguity, multiple priors, and conservative non-Bayesian rules

Several recent papers recast minimal updating in the presence of ambiguity or multiple priors. One direction is conservative updating under ambiguous signals. Han, Jaffray, and Massari propose the conditional maximum likelihood rule (CML) for a max-min expected-utility decision maker with a simple prior set \(\mathcal P\subseteq \Delta(S\times\Theta)\), where all priors share the same marginal \(\mu\) on the payoff-relevant state space \(S\) but differ in the signal interpretation. After observing \(\theta\), CML sets
$$
\mu_{\theta}(s) = \frac{\mu(s)\max_{t\in T} c^t(\theta\mid s)} {\sum_{s'\in S}\mu(s')\max_{t\in T} c^t(\theta\mid s')}.
$$
The paper emphasizes that CML is not a minimum-distance rule and can produce posteriors outside the full Bayesian posterior set \(\mathcal P\mid \theta\). Its conservatism lies elsewhere: it yields a singleton posterior, avoids dilation, satisfies the axiom of Increased Sensitivity after Updating, and is divisible for independent signals [2012.13650].

A different ambiguity-sensitive rule is Robust Maximum Likelihood (RML) updating. Here the decision maker has a benchmark prior \(\pi\) and a set of plausible priors \(\mathbb N\). Upon observing \(A\), RML first restricts to
$$
\mathbb N_A = \left\{ \pi'\in\mathbb N \mid \pi'(A)=\max_{\pi''\in\mathbb N}\pi''(A) \right\},
$$
and then chooses the posterior in \(B(\mathbb N_A)\) minimizing KL divergence from the Bayesian posterior of the benchmark prior:
$$
\pi_A=\underset{\pi_A'\in B(\mathbb{N}_A)}{\arg \min} D_{\text{KL}(\pi(\cdot|A)\: ||\: \pi_A').
$$
The paper is explicit that this is not a pure minimum-updating doctrine. The primary step is evidence-driven prior revision by maximum likelihood; the minimal-change component is a KL tie-break among likelihood-maximizing candidates. Relative to classical minimal revision, RML is therefore a hybrid of prior selection and constrained posterior conservatism [2504.17151].

These ambiguity-sensitive models indicate that minimal updating need not mean preservation of the raw prior. In CML, the central desideratum is avoidance of dilation and preservation of sensitivity after conditioning. In RML, closeness is measured not to the prior itself but to the Bayesian posterior of a benchmark prior, and only after a prior-selection step. A plausible implication is that in ambiguous environments the minimum updating principle often survives only as a secondary criterion subordinate to other structural desiderata, such as maximum likelihood, dynamic consistency, or non-dilation.

## 6. Structural conditions, failures, and methodological implications

The literature repeatedly shows that minimum updating is not self-justifying. One major limitation concerns representation of evidence. Halpern and collaborators argue that the right question is not whether a rule minimizes change in a naive space, but whether it reproduces conditioning in the appropriate sophisticated space that encodes the observation protocol. CAR and generalized CAR delimit when ordinary conditioning and Jeffrey conditioning are legitimate naive-space update rules; no analogous general justification exists for MRE [1407.7183].

A second limitation concerns sufficient statistics for sequential updating. In objective Bayesian inference with continuous parameters and experiment-dependent noninformative priors, ordinary sequential updating may be inconsistent. Kass and Wasserman show that if experiments \(A\) and \(B\) induce different Fisher informations \(h_A(\theta)\) and \(h_B(\theta)\), then carrying forward only the posterior density can produce order dependence. Their revised method updates the cumulative likelihood and the cumulative Fisher information, using
$$
h_C(\theta)=h_A(\theta)+h_B(\theta),
$$
and reconstructs the posterior as
$$
p(\theta\mid x_A,x_B)\propto [L_A(\theta)L_B(\theta)]\,f\!\bigl(h_A(\theta)+h_B(\theta)\bigr),
$$
with \(f(h)=|h|^{1/2}\) for Jeffreys prior. In this setting, posterior-level minimal revision is the wrong criterion; the relevant invariant is the richer informational state consisting of likelihood and Fisher information [1308.2791].

A third limitation is semantic. Smets’ survey, the database-revision literature, and the ambiguity papers all deny that a single syntactic observation determines a unique update. In belief functions, the same event \(A\) can warrant transfer by intersection, proportional renormalization, redistribution to total ignorance, specialization, or imaging. In database revision, minimality is defined only after partitioning the knowledge base into immutable and updatable parts. In ambiguity models, a single posterior may be selected not because it is nearest to the prior, but because it is non-dilating, likelihood-maximizing, or sharply calibrated [1303.5752] [1407.3512] [2012.13650].

The contemporary significance of the minimum updating principle lies precisely in this plurality. It remains a unifying normative intuition, but not a universally executable rule. Across the cited work, the principle is defensible only when the object to be preserved is specified in advance: conditional structure under Bayes, prior joint structure under maximum entropy, product-of-marginals structure under minimum information gain, inclusion-minimal base facts under knowledge-base dynamics, or benchmark-posterior proximity under ambiguity-sensitive updating. The strongest general lesson is therefore methodological rather than algorithmic: before applying any “minimal change” update, one must identify the correct state representation, the active constraints, and the structural invariants that the update is meant to preserve.

Source: https://www.emergentmind.com/topics/minimum-updating-principle