---
title: Personal Decision Theory in Causal Decision Making
url: https://www.emergentmind.com/topics/personal-decision-theory
type: topic
---

# Personal Decision Theory in Causal Decision Making

Personal decision theory concerns decision procedures that target the individual rather than a population average, and, in recent causal formulations, it instructs an agent to maximize a subjective model of that agent’s own counterfactual utility [2208.09558] [2606.29911]. Across the cited work, the term appears in two closely related ways. One usage focuses on personalized decision making in the potential-outcome tradition, where the relevant target is the Individual Treatment Effect (ITE) rather than the Conditional Average Treatment Effect (CATE). A second usage is explicitly decision-theoretic and causal, embedding agents, policies, and counterfactuals in nonparametric structural equation models (NPSEMs) or mechanised causal graphs, and then comparing Evidential Decision Theory (EDT), Causal Decision Theory (CDT), Functional Decision Theory (FDT), and Personal Decision Theory (PDT) itself [2307.10987]. More recent work extends the theme toward reflective support, subjective modeling, and hybrid symbolic-language systems for individual choice [2510.04364].

## 1. Core definitions and scope

In personalized decision making, let $U$ denote a population of units, let $X \in \{0,1\}$ be a binary decision, and let $Y(x,u) \in \{0,1\}$ be the potential outcome of unit $u$ under action $x$. Population-based decision making seeks to maximize the Conditional Average Treatment Effect among those who share observed characteristics $C(u)$:
$$
\mathrm{CATE}(c) \coloneqq E[Y(1,U')-Y(0,U') \mid C(U')=c].
$$
A rule based on $\mathrm{CATE}(c)$ treats every member of the stratum $\{u' : C(u')=c\}$ identically. Personalized decision making instead seeks to optimize the Individual Treatment Effect of unit $u$,
$$
\mathrm{ITE}(u) \coloneqq Y(1,u)-Y(0,u),
$$
and the ideal personalized rule chooses $X=1$ exactly when $\mathrm{ITE}(u)>0$ [2208.09558].

In the NPSEM formulation, each agent is identified with a draw $\epsilon$ of exogenous factors, and PDT instructs the agent to maximize expected counterfactual utility conditional on that $\epsilon$:
$$
A := \arg\max_a E[U(a,\epsilon)\mid \epsilon; M=m].
$$
Under perfect subjective accuracy, this collapses to $A(\epsilon)=\arg\max_a U(a,\epsilon)$ [2606.29911].

| Framework | Target quantity | Canonical decision rule |
|---|---|---|
| Population-based decision making | $\mathrm{CATE}(c)$ | Treat iff $\mathrm{CATE}(C(u))>0$ |
| Personalized decision making | $\mathrm{ITE}(u)$ | Treat iff $\mathrm{ITE}(u)>0$ |
| Personal Decision Theory | $E[U(a,\epsilon)\mid \epsilon;M=m]$ | Choose $A:=\arg\max_a E[U(a,\epsilon)\mid \epsilon;M=m]$ |

A recurring misconception is that “personalized” merely means “conditioned on covariates.” The distinction drawn in this literature is sharper: CATE-based policies remain stratum-level rules, whereas personalized rules aim at latent individual response or individual counterfactual utility. Because each unit experiences only one action, the relevant individualized quantities are counterfactual and require causal assumptions, external information, or bounding arguments rather than direct observation [2208.09558].

## 2. Causal and structural foundations

The potential-outcome presentation uses a simplified causal graph with pre-treatment covariates $C$, decision $X$, and outcome $Y$, where $C \to X \to Y$ and $C$ also points to $Y$. Under the do-operator, one may write
$$
P(Y \mid do(X=x), C=c) = P(Y(x)=1 \mid C=c).
$$
A personalized decision rule $\delta$ assigns to each unit $u$ an action $\delta(u)\in\{0,1\}$ based on all available information, including $C(u)$ and external data. If utility is set to $U(u,x)=Y(x,u)$, then the personalized optimum is
$$
\delta^*(u)=\arg\max_{x\in\{0,1\}} E[U(u,x)\mid \text{information about }u].
$$
When $\mathrm{ITE}(u)$ is known exactly, $\delta^*(u)=1$ if $Y(1,u)-Y(0,u)>0$ [2208.09558].

The NPSEM framework generalizes this language. For endogenous variables $V_1,\ldots,V_k$, an NPSEM specifies independent exogenous variables $\epsilon=(\epsilon_1,\ldots,\epsilon_k)$ with joint density $p(\epsilon)=\prod_i p(\epsilon_i)$ and functions
$$
V_i := f_i(Pa_i,\epsilon_i).
$$
Factual variables $V_i(\epsilon)$ are generated by drawing $\epsilon \sim p(\epsilon)$ and computing the structural equations. Counterfactuals are defined by replacing one equation, for example setting $A:=a$, while leaving all others unchanged, yielding variables denoted $V_i(a)$. Consistency states that if factually $A(\epsilon)=a$, then $V_i(a,\epsilon)=V_i(\epsilon)$ for any other variable $V_i$ [2606.29911].

Mechanised causal Bayesian networks refine this by adding mechanism-level nodes $\tilde V$ so that each object-level variable $V$ has exactly one parent $\tilde V$, whose value encodes the conditional distribution $Pr(V\mid Pa_V)$. In this representation, $\tilde D$ is the decision-rule variable and $U$ is a distinguished utility node. This permits a uniform characterization of conditioning on a realized decision, intervening on a realized decision, and intervening on the decision rule itself [2307.10987].

## 3. Personalized treatment choice and data fusion

Mueller and Pearl formulate the personalized treatment problem with experimental and observational quantities defined on each covariate stratum $c$:
$$
P_t(y\mid c)=P(Y=1\mid do(X=1),C=c), \quad
P_c(y\mid c)=P(Y=1\mid do(X=0),C=c),
$$
$$
P(y\mid t,c)=P(Y=1\mid X=1,C=c), \quad
P(y\mid c,c)=P(Y=1\mid X=0,C=c),
$$
and
$$
\pi(c)=P(X=1\mid C=c).
$$
The central individualized target is
$$
P_{\mathrm{benefit}}(c)\coloneqq P[Y(1,u)=1 \wedge Y(0,u)=0 \mid C(u)=c],
$$
the probability that an individual in stratum $c$ would survive if treated and die if untreated. Its counterpart is
$$
P_{\mathrm{harm}}(c)=P[Y(1,u)=0 \wedge Y(0,u)=1 \mid C(u)=c].
$$
A personalized rule treats exactly those units for which $P_{\mathrm{benefit}}(C(u))>P_{\mathrm{harm}}(C(u))$ [2208.09558].

Randomized controlled trials identify only average differences such as $P_t(y\mid c)-P_c(y\mid c)=\mathrm{CATE}(c)$. Observational data carry confounding biases, but they also encode free-choice mechanisms and therefore information about the joint distribution of $(Y(1),Y(0),X)$. Under consistency,
$$
Y = X\cdot Y(1) + (1-X)\cdot Y(0),
$$
and no selection bias in either study, Tian and Pearl derive tight bounds on $P_{\mathrm{benefit}}(c)$:
$$
\max\left\{
0,\;
P_t(y\mid c)-P_c(y\mid c),\;
P(y\mid t,c)-P_c(y\mid c),\;
P_t(y\mid c)-P(y\mid c,c)
\right\}
\le P_{\mathrm{benefit}}(c)
$$
and
$$
P_{\mathrm{benefit}}(c)\le
\min\left\{
P_t(y\mid c),\;
1-P_c(y\mid c),\;
\pi(c)P(y\mid t,c)+[1-\pi(c)][1-P(y\mid c,c)],
\right.
$$
$$
\left.
[P_t(y\mid c)-P_c(y\mid c)] + \pi(c)[1-P(y\mid t,c)] + [1-\pi(c)]P(y\mid c,c)
\right\}.
$$
When the lower and upper bounds coincide, $P_{\mathrm{benefit}}(c)$ is point-identified. The same exposition gives the looser RCT-only bounds
$$
\max\{0, P_t(y\mid c)-P_c(y\mid c)\}\le P_{\mathrm{benefit}}(c)\le \min\{P_t(y\mid c),1-P_c(y\mid c)\},
$$
and the observational-only bounds
$$
0\le P_{\mathrm{benefit}}(c)\le \pi(c)P(y\mid t,c)+[1-\pi(c)][1-P(y\mid c,c)].
$$
The intended significance is that free-choice observational patterns can sharpen counterfactual probabilities rather than merely introducing nuisance bias [2208.09558].

The motivating drug-trial example uses sex as the only observed covariate. In the randomized trial, females have $P_t(y\mid female)=0.489$, $P_c(y\mid female)=0.210$, and $\mathrm{CATE}(female)=0.279$, while males have $P_t(y\mid male)=0.490$, $P_c(y\mid male)=0.210$, and $\mathrm{CATE}(male)=0.280$. The observational study reports that for females, $70\%$ chose the drug, $P(y\mid t,female)=378/1400=0.270$, $P(y\mid c,female)=420/600=0.700$, and $\pi(female)=0.70$; for males, $70\%$ chose the drug, $P(y\mid t,male)=980/1400=0.700$, $P(y\mid c,male)=420/600=0.700$, and $\pi(male)=0.70$. Applying the Tian–Pearl bounds yields point identification:
$$
P_{\mathrm{benefit}}(female)=0.279,\quad P_{\mathrm{harm}}(female)=0.000,
$$
$$
P_{\mathrm{benefit}}(male)=0.490,\quad P_{\mathrm{harm}}(male)=0.210.
$$
Thus 28% of women individually benefit and none are harmed, whereas 49% of men individually benefit and 21% are harmed, despite nearly identical RCT averages [2208.09558].

## 4. Normative PDT in the NPSEM framework

Sjölander’s formulation adds two endogenous nodes to an NPSEM: $D \in \{\mathrm{EDT},\mathrm{CDT},\mathrm{PDT},\ldots\}$, recording which decision theory the agent uses, and $M$, encoding the agent’s subjective model of the world. The act $A$ is then an endogenous variable determined by $D$ and $M$. EDT is defined by
$$
A := \arg\max_a E[U \mid A=a; M=m],
$$
CDT by
$$
A := \arg\max_a E[U(a); M=m],
$$
and PDT by
$$
A := \arg\max_a E[U(a,\epsilon)\mid \epsilon; M=m].
$$
The distinctive feature of PDT is that it conditions on the exogenous realization identifying the particular agent, rather than maximizing a population-level expectation [2606.29911].

The framework also introduces an objective performance metric based on population interventions. If a policy-maker could enforce “everyone uses decision theory $d$,” then $U(d)$ denotes the counterfactual utility when $D:=d$ for every agent, and performance is
$$
\mathrm{Perf}(d)\coloneqq E[U(d)].
$$
This turns the evaluation of decision theories into a causal question about interventions on the decision-theory node itself [2606.29911].

Three results organize the comparison. Assumption 1 requires correct subjective models: under EDT, the agent’s subjective $E(U\mid A=a;M=\mathrm{EDT})$ equals the true $E(U\mid A=a)$; under CDT, subjective $E(U(a);M=\mathrm{CDT})$ equals true $E(U(a))$; under PDT, $E(U(a,\epsilon)\mid \epsilon;M=\mathrm{PDT})=U(a,\epsilon)$ for all $\epsilon$. Assumption 2 requires no direct effect of $D$ on $U$, namely $U(a,d)=U(a)$. Under Assumption 1, EDT and CDT choose the same constant act for all $\epsilon$. Under Assumptions 1–2, CDT weakly dominates any decision rule whose implied act is constant across $\epsilon$. Under Assumptions 1–2, PDT is globally optimal:
$$
\mathrm{Perf}(\mathrm{PDT}) \ge \mathrm{Perf}(d)\quad \text{for every decision-rule } d.
$$
The proof sketch is pointwise: $\mathrm{Perf}(d)=E[U(A(d),\epsilon)]$, and PDT chooses $A$ maximizing $U(a,\epsilon)$ for each $\epsilon$, so integration preserves the inequality [2606.29911].

These theorems delimit, rather than eliminate, controversy. They establish PDT’s optimality only under specific assumptions about subjective accuracy and about the absence of a direct pathway from decision theory to utility. The framework therefore clarifies which disputes are substantive disagreements about causal structure and which are merely terminological disagreements about what counts as the relevant expectation.

## 5. Canonical paradoxes, mechanised taxonomy, and controversy

The smoking–lesion problem is used to display the divergence between evidential, causal, and personal criteria. In Sjölander’s NPSEM, $G$ denotes genetic lesion presence, $A$ smoking or not, $Y$ cancer or not, and $U$ utility, with equations
$$
G:=f_G(\epsilon_G),\quad
A:=f_A(G,\epsilon_A),\quad
Y:=f_Y(G,A,\epsilon_Y),\quad
U:=f_U(A,Y,\epsilon_U).
$$
Here EDT chooses $A=\arg\max_a E[U\mid A=a]$, CDT chooses $A=\arg\max_a E[U(a)]$, and PDT chooses $A(\epsilon)=\arg\max_a U(a,\epsilon)$. In many versions of the problem, $E(U\mid A=1)<E(U\mid A=0)$ but $E(U(1))>E(U(0))$, so EDT does not smoke, CDT does smoke, and PDT tailors the act to the particular $\epsilon$ [2606.29911].

Mechanised causal graphs provide a broader taxonomy. Two axes generate six procedures: method—Evidential, Causal, Functional—and updatefulness—Updateful or Updateless. Standard EDT is updateful evidential choice,
$$
d^*_{\mathrm{EDT}} \in \arg\max_d E[U\mid D=d, Obs_D=obs_D].
$$
Standard CDT is updateful causal choice,
$$
d^*_{\mathrm{CDT}} \in \arg\max_d E[U\mid do(D=d), Obs_D=obs_D].
$$
Standard FDT is updateless functional choice,
$$
\tilde d^*_{\mathrm{FDT}} \in \arg\max_{\tilde d} E[U\mid do(\tilde D=\tilde d)].
$$
The mechanised formalism emphasizes that the theories differ not only in what they optimize, but also in whether they condition on the realized action, intervene on the action, or intervene on the mechanism generating the action [2307.10987].

Newcomb’s problem exposes the importance of direct causal effects of the decision theory. In Sjölander’s NPSEM, variables include agent characteristics $G$, decision theory $D$, model $M$, predictor’s prediction $Y$, act $A$, and payoff $U=x\cdot A + z\cdot(1-Y)$. Under Assumption 1, and provided the predictor is sufficiently accurate so that $E(U\mid A=0)>E(U\mid A=1)$, the analysis yields
$$
A(\mathrm{CDT})=A(\mathrm{PDT})=1,\qquad A(\mathrm{EDT})=0.
$$
Yet the same discussion notes that $D \to Y \to U$ gives a direct effect of $D$ on $U$, so the global optimality theorem for PDT need not apply [2606.29911]. In the mechanised-graph formulation, by contrast, FDT is represented as an intervention on $\tilde D$ in a logical-causal model, which reproduces one-boxing in Newcomb and cooperation in Twin Prisoner’s Dilemma when predictors or twins depend on the same mechanism [2307.10987]. The resulting controversy is not erased by the formalism; rather, the formalism makes explicit which arrows, observations, and intervention targets produce each verdict.

## 6. Reflective, predictive, and adjacent extensions

A more human-centered line of work connects personal decision theory to metacognition rather than only to counterfactual choice rules. In this usage, personal decision theory emphasizes that people rarely make choices in a purely rational, utility-maximizing way; bounded cognitive resources, heuristics, and affective biases steer decision making toward solution-driven shortcuts. “Pre-Decision Reflection” shifts assistance from “here’s what you should do” to “here’s how you are thinking about it.” PROBE operationalizes this with two orthogonal metrics over seven reflective aspects—Belief, Awareness of Difficulties, Experience, Feeling, Intention, Insight, and Alternative Perspective. Breadth is
$$
B(r)=\sum_{c=1}^{C}\mathbf{1}[n_c(r)>0], \qquad C=7,
$$
and depth is
$$
D(r)=\frac{E(r)}{N(r)}\times 100\%.
$$
In the reported sample, breadth ranged from $B=1$ to $B=6$ with mean $\mu_B=3.2$ and $SD_B=1.34$, while 80% of participants had $D(r)<50\%$. Reliability was established with Fleiss’s $\kappa=0.69$ on a triple-coded subset and Cohen’s $\kappa=0.79$ on a double-coded subset [2510.04364].

Another extension models individualized choice probabilistically. In the Quantum Decision Theory account, the probability of choosing prospect $\pi_i$ is
$$
P(\pi_i)=f(\pi_i)+q(\pi_i),
$$
where $f(\pi_i)$ is the utility factor and $q(\pi_i)$ is the attraction factor. For binary choice between a gamble $A$ and a sure option $B$, the classical component is
$$
f_A=\frac{1}{1+e^{\phi(U_B-U_A)}},
$$
while the attraction component takes the form
$$
q_A=\min(f_A,f_B)\cos\Delta,\qquad q_B=-q_A.
$$
The interference phase is parameterized through framing, time pressure, memory, and need. In the reported experiments, the full QDT model achieved $81.2\%\pm 0.8\%$ on Dataset 1 and $72.4\%\pm 0.6\%$ on Dataset 2, outperforming the stated CPT and machine-learning baselines, with paired $t$-test significance at $p<0.01$ [2101.05851].

A separate computational line combines symbolic utility discovery with language-based personalization. ATHENA first discovers group-level symbolic utility functions
$$
U_j^{(g)}(X)=f_g^*(X,j;\theta_g^*)
$$
for each alternative $j$, and then initializes and optimizes an individual textual template
$$
\mathcal P_i^0 \sim \phi(\cdot \mid f_g^*, i, \mathcal C), \qquad
\mathcal P_i^{t+1}=\mathcal P_i^t-\eta \nabla_{\mathcal P_i^t}\mathcal L_i(\mathcal P_i^t;D_i).
$$
Final prediction is generated from the adapted template together with the utility-guided context. On Swissmetro, ATHENA achieved $\mathrm{Acc}=0.813$, $\mathrm{F1}=0.766$, $\mathrm{AUC}=0.883$, and $\mathrm{CE}=1.086$; on Vaccine, $\mathrm{Acc}=0.735$, $\mathrm{F1}=0.716$, $\mathrm{AUC}=0.870$, and $\mathrm{CE}=0.755$. The ablation study reported that removing symbolic discovery or semantic adaptation drops accuracy by 10–15 points [2511.02194].

Related individual-level models broaden the surrounding landscape. The justifiability model posits a true preference $\succ^*$ together with a family $\mathcal J$ of justifiable total orders; for any menu $A$, the justifiable candidates are
$$
M(A):=\bigcup_{\succ_j\in\mathcal J}\arg\max(A,\succ_j),
$$
and choice is
$$
c(A)=\arg\max(M(A),\succ^*).
$$
This captures decision making constrained by morality, rationality, or other virtues [2003.06844]. Hope-and-prepare preferences instead define an incomplete ordering under uncertainty in which $f\succ g$ only when both a pessimistic minimum over one set of priors $C$ and an optimistic maximum over another set of priors $D$ favor $f$ over $g$:
$$
\min_{p\in C}\int_S u(f(s))\,dp(s)>\min_{p\in C}\int_S u(g(s))\,dp(s),
$$
and
$$
\max_{p\in D}\int_S u(f(s))\,dp(s)>\max_{p\in D}\int_S u(g(s))\,dp(s).
$$
This model is described as “hopes for the best” while “prepares for the worst” [2406.11166].

Taken together, these works suggest a common research program: the relevant object is not merely expected utility averaged over a population, but decision quality indexed to the particular agent, that agent’s latent counterfactuals, subjective model, justificatory constraints, reflective profile, or semantic context. The main unresolved disputes then concern which causal graph is appropriate, what information may legitimately condition the choice rule, whether observational regularities should refine individualized counterfactuals, and how far individual decision support should optimize outcomes versus cultivate self-awareness.

Source: https://www.emergentmind.com/topics/personal-decision-theory