Papers
Topics
Authors
Recent
Search
2000 character limit reached

Personal Decision Theory in Causal Decision Making

Updated 6 July 2026
  • Personal Decision Theory is a framework that optimizes individual decisions by maximizing counterfactual utility using both potential outcome and causal intervention methods.
  • It distinguishes between population-level averages like CATE and individualized measures such as ITE, enabling tailored treatment choices.
  • Advanced implementations leverage NPSEM, mechanised causal graphs, and hybrid symbolic-language systems to refine personalized decision-making.

Personal decision theory concerns decision procedures that target the individual rather than a population average, and, in recent causal formulations, it instructs an agent to maximize a subjective model of that agent’s own counterfactual utility (Mueller et al., 2022, Sjölander, 29 Jun 2026). Across the cited work, the term appears in two closely related ways. One usage focuses on personalized decision making in the potential-outcome tradition, where the relevant target is the Individual Treatment Effect (ITE) rather than the Conditional Average Treatment Effect (CATE). A second usage is explicitly decision-theoretic and causal, embedding agents, policies, and counterfactuals in nonparametric structural equation models (NPSEMs) or mechanised causal graphs, and then comparing Evidential Decision Theory (EDT), Causal Decision Theory (CDT), Functional Decision Theory (FDT), and Personal Decision Theory (PDT) itself (MacDermott et al., 2023). More recent work extends the theme toward reflective support, subjective modeling, and hybrid symbolic-language systems for individual choice (Tarvirdians et al., 5 Oct 2025).

1. Core definitions and scope

In personalized decision making, let UU denote a population of units, let X{0,1}X \in \{0,1\} be a binary decision, and let Y(x,u){0,1}Y(x,u) \in \{0,1\} be the potential outcome of unit uu under action xx. Population-based decision making seeks to maximize the Conditional Average Treatment Effect among those who share observed characteristics C(u)C(u):

CATE(c)E[Y(1,U)Y(0,U)C(U)=c].\mathrm{CATE}(c) \coloneqq E[Y(1,U')-Y(0,U') \mid C(U')=c].

A rule based on CATE(c)\mathrm{CATE}(c) treats every member of the stratum {u:C(u)=c}\{u' : C(u')=c\} identically. Personalized decision making instead seeks to optimize the Individual Treatment Effect of unit uu,

X{0,1}X \in \{0,1\}0

and the ideal personalized rule chooses X{0,1}X \in \{0,1\}1 exactly when X{0,1}X \in \{0,1\}2 (Mueller et al., 2022).

In the NPSEM formulation, each agent is identified with a draw X{0,1}X \in \{0,1\}3 of exogenous factors, and PDT instructs the agent to maximize expected counterfactual utility conditional on that X{0,1}X \in \{0,1\}4:

X{0,1}X \in \{0,1\}5

Under perfect subjective accuracy, this collapses to X{0,1}X \in \{0,1\}6 (Sjölander, 29 Jun 2026).

Framework Target quantity Canonical decision rule
Population-based decision making X{0,1}X \in \{0,1\}7 Treat iff X{0,1}X \in \{0,1\}8
Personalized decision making X{0,1}X \in \{0,1\}9 Treat iff Y(x,u){0,1}Y(x,u) \in \{0,1\}0
Personal Decision Theory Y(x,u){0,1}Y(x,u) \in \{0,1\}1 Choose Y(x,u){0,1}Y(x,u) \in \{0,1\}2

A recurring misconception is that “personalized” merely means “conditioned on covariates.” The distinction drawn in this literature is sharper: CATE-based policies remain stratum-level rules, whereas personalized rules aim at latent individual response or individual counterfactual utility. Because each unit experiences only one action, the relevant individualized quantities are counterfactual and require causal assumptions, external information, or bounding arguments rather than direct observation (Mueller et al., 2022).

2. Causal and structural foundations

The potential-outcome presentation uses a simplified causal graph with pre-treatment covariates Y(x,u){0,1}Y(x,u) \in \{0,1\}3, decision Y(x,u){0,1}Y(x,u) \in \{0,1\}4, and outcome Y(x,u){0,1}Y(x,u) \in \{0,1\}5, where Y(x,u){0,1}Y(x,u) \in \{0,1\}6 and Y(x,u){0,1}Y(x,u) \in \{0,1\}7 also points to Y(x,u){0,1}Y(x,u) \in \{0,1\}8. Under the do-operator, one may write

Y(x,u){0,1}Y(x,u) \in \{0,1\}9

A personalized decision rule uu0 assigns to each unit uu1 an action uu2 based on all available information, including uu3 and external data. If utility is set to uu4, then the personalized optimum is

uu5

When uu6 is known exactly, uu7 if uu8 (Mueller et al., 2022).

The NPSEM framework generalizes this language. For endogenous variables uu9, an NPSEM specifies independent exogenous variables xx0 with joint density xx1 and functions

xx2

Factual variables xx3 are generated by drawing xx4 and computing the structural equations. Counterfactuals are defined by replacing one equation, for example setting xx5, while leaving all others unchanged, yielding variables denoted xx6. Consistency states that if factually xx7, then xx8 for any other variable xx9 (Sjölander, 29 Jun 2026).

Mechanised causal Bayesian networks refine this by adding mechanism-level nodes C(u)C(u)0 so that each object-level variable C(u)C(u)1 has exactly one parent C(u)C(u)2, whose value encodes the conditional distribution C(u)C(u)3. In this representation, C(u)C(u)4 is the decision-rule variable and C(u)C(u)5 is a distinguished utility node. This permits a uniform characterization of conditioning on a realized decision, intervening on a realized decision, and intervening on the decision rule itself (MacDermott et al., 2023).

3. Personalized treatment choice and data fusion

Mueller and Pearl formulate the personalized treatment problem with experimental and observational quantities defined on each covariate stratum C(u)C(u)6:

C(u)C(u)7

C(u)C(u)8

and

C(u)C(u)9

The central individualized target is

CATE(c)E[Y(1,U)Y(0,U)C(U)=c].\mathrm{CATE}(c) \coloneqq E[Y(1,U')-Y(0,U') \mid C(U')=c].0

the probability that an individual in stratum CATE(c)E[Y(1,U)Y(0,U)C(U)=c].\mathrm{CATE}(c) \coloneqq E[Y(1,U')-Y(0,U') \mid C(U')=c].1 would survive if treated and die if untreated. Its counterpart is

CATE(c)E[Y(1,U)Y(0,U)C(U)=c].\mathrm{CATE}(c) \coloneqq E[Y(1,U')-Y(0,U') \mid C(U')=c].2

A personalized rule treats exactly those units for which CATE(c)E[Y(1,U)Y(0,U)C(U)=c].\mathrm{CATE}(c) \coloneqq E[Y(1,U')-Y(0,U') \mid C(U')=c].3 (Mueller et al., 2022).

Randomized controlled trials identify only average differences such as CATE(c)E[Y(1,U)Y(0,U)C(U)=c].\mathrm{CATE}(c) \coloneqq E[Y(1,U')-Y(0,U') \mid C(U')=c].4. Observational data carry confounding biases, but they also encode free-choice mechanisms and therefore information about the joint distribution of CATE(c)E[Y(1,U)Y(0,U)C(U)=c].\mathrm{CATE}(c) \coloneqq E[Y(1,U')-Y(0,U') \mid C(U')=c].5. Under consistency,

CATE(c)E[Y(1,U)Y(0,U)C(U)=c].\mathrm{CATE}(c) \coloneqq E[Y(1,U')-Y(0,U') \mid C(U')=c].6

and no selection bias in either study, Tian and Pearl derive tight bounds on CATE(c)E[Y(1,U)Y(0,U)C(U)=c].\mathrm{CATE}(c) \coloneqq E[Y(1,U')-Y(0,U') \mid C(U')=c].7:

CATE(c)E[Y(1,U)Y(0,U)C(U)=c].\mathrm{CATE}(c) \coloneqq E[Y(1,U')-Y(0,U') \mid C(U')=c].8

and

CATE(c)E[Y(1,U)Y(0,U)C(U)=c].\mathrm{CATE}(c) \coloneqq E[Y(1,U')-Y(0,U') \mid C(U')=c].9

CATE(c)\mathrm{CATE}(c)0

When the lower and upper bounds coincide, CATE(c)\mathrm{CATE}(c)1 is point-identified. The same exposition gives the looser RCT-only bounds

CATE(c)\mathrm{CATE}(c)2

and the observational-only bounds

CATE(c)\mathrm{CATE}(c)3

The intended significance is that free-choice observational patterns can sharpen counterfactual probabilities rather than merely introducing nuisance bias (Mueller et al., 2022).

The motivating drug-trial example uses sex as the only observed covariate. In the randomized trial, females have CATE(c)\mathrm{CATE}(c)4, CATE(c)\mathrm{CATE}(c)5, and CATE(c)\mathrm{CATE}(c)6, while males have CATE(c)\mathrm{CATE}(c)7, CATE(c)\mathrm{CATE}(c)8, and CATE(c)\mathrm{CATE}(c)9. The observational study reports that for females, {u:C(u)=c}\{u' : C(u')=c\}0 chose the drug, {u:C(u)=c}\{u' : C(u')=c\}1, {u:C(u)=c}\{u' : C(u')=c\}2, and {u:C(u)=c}\{u' : C(u')=c\}3; for males, {u:C(u)=c}\{u' : C(u')=c\}4 chose the drug, {u:C(u)=c}\{u' : C(u')=c\}5, {u:C(u)=c}\{u' : C(u')=c\}6, and {u:C(u)=c}\{u' : C(u')=c\}7. Applying the Tian–Pearl bounds yields point identification:

{u:C(u)=c}\{u' : C(u')=c\}8

{u:C(u)=c}\{u' : C(u')=c\}9

Thus 28% of women individually benefit and none are harmed, whereas 49% of men individually benefit and 21% are harmed, despite nearly identical RCT averages (Mueller et al., 2022).

4. Normative PDT in the NPSEM framework

Sjölander’s formulation adds two endogenous nodes to an NPSEM: uu0, recording which decision theory the agent uses, and uu1, encoding the agent’s subjective model of the world. The act uu2 is then an endogenous variable determined by uu3 and uu4. EDT is defined by

uu5

CDT by

uu6

and PDT by

uu7

The distinctive feature of PDT is that it conditions on the exogenous realization identifying the particular agent, rather than maximizing a population-level expectation (Sjölander, 29 Jun 2026).

The framework also introduces an objective performance metric based on population interventions. If a policy-maker could enforce “everyone uses decision theory uu8,” then uu9 denotes the counterfactual utility when X{0,1}X \in \{0,1\}00 for every agent, and performance is

X{0,1}X \in \{0,1\}01

This turns the evaluation of decision theories into a causal question about interventions on the decision-theory node itself (Sjölander, 29 Jun 2026).

Three results organize the comparison. Assumption 1 requires correct subjective models: under EDT, the agent’s subjective X{0,1}X \in \{0,1\}02 equals the true X{0,1}X \in \{0,1\}03; under CDT, subjective X{0,1}X \in \{0,1\}04 equals true X{0,1}X \in \{0,1\}05; under PDT, X{0,1}X \in \{0,1\}06 for all X{0,1}X \in \{0,1\}07. Assumption 2 requires no direct effect of X{0,1}X \in \{0,1\}08 on X{0,1}X \in \{0,1\}09, namely X{0,1}X \in \{0,1\}10. Under Assumption 1, EDT and CDT choose the same constant act for all X{0,1}X \in \{0,1\}11. Under Assumptions 1–2, CDT weakly dominates any decision rule whose implied act is constant across X{0,1}X \in \{0,1\}12. Under Assumptions 1–2, PDT is globally optimal:

X{0,1}X \in \{0,1\}13

The proof sketch is pointwise: X{0,1}X \in \{0,1\}14, and PDT chooses X{0,1}X \in \{0,1\}15 maximizing X{0,1}X \in \{0,1\}16 for each X{0,1}X \in \{0,1\}17, so integration preserves the inequality (Sjölander, 29 Jun 2026).

These theorems delimit, rather than eliminate, controversy. They establish PDT’s optimality only under specific assumptions about subjective accuracy and about the absence of a direct pathway from decision theory to utility. The framework therefore clarifies which disputes are substantive disagreements about causal structure and which are merely terminological disagreements about what counts as the relevant expectation.

5. Canonical paradoxes, mechanised taxonomy, and controversy

The smoking–lesion problem is used to display the divergence between evidential, causal, and personal criteria. In Sjölander’s NPSEM, X{0,1}X \in \{0,1\}18 denotes genetic lesion presence, X{0,1}X \in \{0,1\}19 smoking or not, X{0,1}X \in \{0,1\}20 cancer or not, and X{0,1}X \in \{0,1\}21 utility, with equations

X{0,1}X \in \{0,1\}22

Here EDT chooses X{0,1}X \in \{0,1\}23, CDT chooses X{0,1}X \in \{0,1\}24, and PDT chooses X{0,1}X \in \{0,1\}25. In many versions of the problem, X{0,1}X \in \{0,1\}26 but X{0,1}X \in \{0,1\}27, so EDT does not smoke, CDT does smoke, and PDT tailors the act to the particular X{0,1}X \in \{0,1\}28 (Sjölander, 29 Jun 2026).

Mechanised causal graphs provide a broader taxonomy. Two axes generate six procedures: method—Evidential, Causal, Functional—and updatefulness—Updateful or Updateless. Standard EDT is updateful evidential choice,

X{0,1}X \in \{0,1\}29

Standard CDT is updateful causal choice,

X{0,1}X \in \{0,1\}30

Standard FDT is updateless functional choice,

X{0,1}X \in \{0,1\}31

The mechanised formalism emphasizes that the theories differ not only in what they optimize, but also in whether they condition on the realized action, intervene on the action, or intervene on the mechanism generating the action (MacDermott et al., 2023).

Newcomb’s problem exposes the importance of direct causal effects of the decision theory. In Sjölander’s NPSEM, variables include agent characteristics X{0,1}X \in \{0,1\}32, decision theory X{0,1}X \in \{0,1\}33, model X{0,1}X \in \{0,1\}34, predictor’s prediction X{0,1}X \in \{0,1\}35, act X{0,1}X \in \{0,1\}36, and payoff X{0,1}X \in \{0,1\}37. Under Assumption 1, and provided the predictor is sufficiently accurate so that X{0,1}X \in \{0,1\}38, the analysis yields

X{0,1}X \in \{0,1\}39

Yet the same discussion notes that X{0,1}X \in \{0,1\}40 gives a direct effect of X{0,1}X \in \{0,1\}41 on X{0,1}X \in \{0,1\}42, so the global optimality theorem for PDT need not apply (Sjölander, 29 Jun 2026). In the mechanised-graph formulation, by contrast, FDT is represented as an intervention on X{0,1}X \in \{0,1\}43 in a logical-causal model, which reproduces one-boxing in Newcomb and cooperation in Twin Prisoner’s Dilemma when predictors or twins depend on the same mechanism (MacDermott et al., 2023). The resulting controversy is not erased by the formalism; rather, the formalism makes explicit which arrows, observations, and intervention targets produce each verdict.

6. Reflective, predictive, and adjacent extensions

A more human-centered line of work connects personal decision theory to metacognition rather than only to counterfactual choice rules. In this usage, personal decision theory emphasizes that people rarely make choices in a purely rational, utility-maximizing way; bounded cognitive resources, heuristics, and affective biases steer decision making toward solution-driven shortcuts. “Pre-Decision Reflection” shifts assistance from “here’s what you should do” to “here’s how you are thinking about it.” PROBE operationalizes this with two orthogonal metrics over seven reflective aspects—Belief, Awareness of Difficulties, Experience, Feeling, Intention, Insight, and Alternative Perspective. Breadth is

X{0,1}X \in \{0,1\}44

and depth is

X{0,1}X \in \{0,1\}45

In the reported sample, breadth ranged from X{0,1}X \in \{0,1\}46 to X{0,1}X \in \{0,1\}47 with mean X{0,1}X \in \{0,1\}48 and X{0,1}X \in \{0,1\}49, while 80% of participants had X{0,1}X \in \{0,1\}50. Reliability was established with Fleiss’s X{0,1}X \in \{0,1\}51 on a triple-coded subset and Cohen’s X{0,1}X \in \{0,1\}52 on a double-coded subset (Tarvirdians et al., 5 Oct 2025).

Another extension models individualized choice probabilistically. In the Quantum Decision Theory account, the probability of choosing prospect X{0,1}X \in \{0,1\}53 is

X{0,1}X \in \{0,1\}54

where X{0,1}X \in \{0,1\}55 is the utility factor and X{0,1}X \in \{0,1\}56 is the attraction factor. For binary choice between a gamble X{0,1}X \in \{0,1\}57 and a sure option X{0,1}X \in \{0,1\}58, the classical component is

X{0,1}X \in \{0,1\}59

while the attraction component takes the form

X{0,1}X \in \{0,1\}60

The interference phase is parameterized through framing, time pressure, memory, and need. In the reported experiments, the full QDT model achieved X{0,1}X \in \{0,1\}61 on Dataset 1 and X{0,1}X \in \{0,1\}62 on Dataset 2, outperforming the stated CPT and machine-learning baselines, with paired X{0,1}X \in \{0,1\}63-test significance at X{0,1}X \in \{0,1\}64 (Zhang et al., 2021).

A separate computational line combines symbolic utility discovery with language-based personalization. ATHENA first discovers group-level symbolic utility functions

X{0,1}X \in \{0,1\}65

for each alternative X{0,1}X \in \{0,1\}66, and then initializes and optimizes an individual textual template

X{0,1}X \in \{0,1\}67

Final prediction is generated from the adapted template together with the utility-guided context. On Swissmetro, ATHENA achieved X{0,1}X \in \{0,1\}68, X{0,1}X \in \{0,1\}69, X{0,1}X \in \{0,1\}70, and X{0,1}X \in \{0,1\}71; on Vaccine, X{0,1}X \in \{0,1\}72, X{0,1}X \in \{0,1\}73, X{0,1}X \in \{0,1\}74, and X{0,1}X \in \{0,1\}75. The ablation study reported that removing symbolic discovery or semantic adaptation drops accuracy by 10–15 points (Zhao et al., 4 Nov 2025).

Related individual-level models broaden the surrounding landscape. The justifiability model posits a true preference X{0,1}X \in \{0,1\}76 together with a family X{0,1}X \in \{0,1\}77 of justifiable total orders; for any menu X{0,1}X \in \{0,1\}78, the justifiable candidates are

X{0,1}X \in \{0,1\}79

and choice is

X{0,1}X \in \{0,1\}80

This captures decision making constrained by morality, rationality, or other virtues (Ridout, 2020). Hope-and-prepare preferences instead define an incomplete ordering under uncertainty in which X{0,1}X \in \{0,1\}81 only when both a pessimistic minimum over one set of priors X{0,1}X \in \{0,1\}82 and an optimistic maximum over another set of priors X{0,1}X \in \{0,1\}83 favor X{0,1}X \in \{0,1\}84 over X{0,1}X \in \{0,1\}85:

X{0,1}X \in \{0,1\}86

and

X{0,1}X \in \{0,1\}87

This model is described as “hopes for the best” while “prepares for the worst” (Bardier et al., 2024).

Taken together, these works suggest a common research program: the relevant object is not merely expected utility averaged over a population, but decision quality indexed to the particular agent, that agent’s latent counterfactuals, subjective model, justificatory constraints, reflective profile, or semantic context. The main unresolved disputes then concern which causal graph is appropriate, what information may legitimately condition the choice rule, whether observational regularities should refine individualized counterfactuals, and how far individual decision support should optimize outcomes versus cultivate self-awareness.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Personal Decision Theory.