Personal Decision Theory in Causal Decision Making
- Personal Decision Theory is a framework that optimizes individual decisions by maximizing counterfactual utility using both potential outcome and causal intervention methods.
- It distinguishes between population-level averages like CATE and individualized measures such as ITE, enabling tailored treatment choices.
- Advanced implementations leverage NPSEM, mechanised causal graphs, and hybrid symbolic-language systems to refine personalized decision-making.
Personal decision theory concerns decision procedures that target the individual rather than a population average, and, in recent causal formulations, it instructs an agent to maximize a subjective model of that agent’s own counterfactual utility (Mueller et al., 2022, Sjölander, 29 Jun 2026). Across the cited work, the term appears in two closely related ways. One usage focuses on personalized decision making in the potential-outcome tradition, where the relevant target is the Individual Treatment Effect (ITE) rather than the Conditional Average Treatment Effect (CATE). A second usage is explicitly decision-theoretic and causal, embedding agents, policies, and counterfactuals in nonparametric structural equation models (NPSEMs) or mechanised causal graphs, and then comparing Evidential Decision Theory (EDT), Causal Decision Theory (CDT), Functional Decision Theory (FDT), and Personal Decision Theory (PDT) itself (MacDermott et al., 2023). More recent work extends the theme toward reflective support, subjective modeling, and hybrid symbolic-language systems for individual choice (Tarvirdians et al., 5 Oct 2025).
1. Core definitions and scope
In personalized decision making, let denote a population of units, let be a binary decision, and let be the potential outcome of unit under action . Population-based decision making seeks to maximize the Conditional Average Treatment Effect among those who share observed characteristics :
A rule based on treats every member of the stratum identically. Personalized decision making instead seeks to optimize the Individual Treatment Effect of unit ,
0
and the ideal personalized rule chooses 1 exactly when 2 (Mueller et al., 2022).
In the NPSEM formulation, each agent is identified with a draw 3 of exogenous factors, and PDT instructs the agent to maximize expected counterfactual utility conditional on that 4:
5
Under perfect subjective accuracy, this collapses to 6 (Sjölander, 29 Jun 2026).
| Framework | Target quantity | Canonical decision rule |
|---|---|---|
| Population-based decision making | 7 | Treat iff 8 |
| Personalized decision making | 9 | Treat iff 0 |
| Personal Decision Theory | 1 | Choose 2 |
A recurring misconception is that “personalized” merely means “conditioned on covariates.” The distinction drawn in this literature is sharper: CATE-based policies remain stratum-level rules, whereas personalized rules aim at latent individual response or individual counterfactual utility. Because each unit experiences only one action, the relevant individualized quantities are counterfactual and require causal assumptions, external information, or bounding arguments rather than direct observation (Mueller et al., 2022).
2. Causal and structural foundations
The potential-outcome presentation uses a simplified causal graph with pre-treatment covariates 3, decision 4, and outcome 5, where 6 and 7 also points to 8. Under the do-operator, one may write
9
A personalized decision rule 0 assigns to each unit 1 an action 2 based on all available information, including 3 and external data. If utility is set to 4, then the personalized optimum is
5
When 6 is known exactly, 7 if 8 (Mueller et al., 2022).
The NPSEM framework generalizes this language. For endogenous variables 9, an NPSEM specifies independent exogenous variables 0 with joint density 1 and functions
2
Factual variables 3 are generated by drawing 4 and computing the structural equations. Counterfactuals are defined by replacing one equation, for example setting 5, while leaving all others unchanged, yielding variables denoted 6. Consistency states that if factually 7, then 8 for any other variable 9 (Sjölander, 29 Jun 2026).
Mechanised causal Bayesian networks refine this by adding mechanism-level nodes 0 so that each object-level variable 1 has exactly one parent 2, whose value encodes the conditional distribution 3. In this representation, 4 is the decision-rule variable and 5 is a distinguished utility node. This permits a uniform characterization of conditioning on a realized decision, intervening on a realized decision, and intervening on the decision rule itself (MacDermott et al., 2023).
3. Personalized treatment choice and data fusion
Mueller and Pearl formulate the personalized treatment problem with experimental and observational quantities defined on each covariate stratum 6:
7
8
and
9
The central individualized target is
0
the probability that an individual in stratum 1 would survive if treated and die if untreated. Its counterpart is
2
A personalized rule treats exactly those units for which 3 (Mueller et al., 2022).
Randomized controlled trials identify only average differences such as 4. Observational data carry confounding biases, but they also encode free-choice mechanisms and therefore information about the joint distribution of 5. Under consistency,
6
and no selection bias in either study, Tian and Pearl derive tight bounds on 7:
8
and
9
0
When the lower and upper bounds coincide, 1 is point-identified. The same exposition gives the looser RCT-only bounds
2
and the observational-only bounds
3
The intended significance is that free-choice observational patterns can sharpen counterfactual probabilities rather than merely introducing nuisance bias (Mueller et al., 2022).
The motivating drug-trial example uses sex as the only observed covariate. In the randomized trial, females have 4, 5, and 6, while males have 7, 8, and 9. The observational study reports that for females, 0 chose the drug, 1, 2, and 3; for males, 4 chose the drug, 5, 6, and 7. Applying the Tian–Pearl bounds yields point identification:
8
9
Thus 28% of women individually benefit and none are harmed, whereas 49% of men individually benefit and 21% are harmed, despite nearly identical RCT averages (Mueller et al., 2022).
4. Normative PDT in the NPSEM framework
Sjölander’s formulation adds two endogenous nodes to an NPSEM: 0, recording which decision theory the agent uses, and 1, encoding the agent’s subjective model of the world. The act 2 is then an endogenous variable determined by 3 and 4. EDT is defined by
5
CDT by
6
and PDT by
7
The distinctive feature of PDT is that it conditions on the exogenous realization identifying the particular agent, rather than maximizing a population-level expectation (Sjölander, 29 Jun 2026).
The framework also introduces an objective performance metric based on population interventions. If a policy-maker could enforce “everyone uses decision theory 8,” then 9 denotes the counterfactual utility when 00 for every agent, and performance is
01
This turns the evaluation of decision theories into a causal question about interventions on the decision-theory node itself (Sjölander, 29 Jun 2026).
Three results organize the comparison. Assumption 1 requires correct subjective models: under EDT, the agent’s subjective 02 equals the true 03; under CDT, subjective 04 equals true 05; under PDT, 06 for all 07. Assumption 2 requires no direct effect of 08 on 09, namely 10. Under Assumption 1, EDT and CDT choose the same constant act for all 11. Under Assumptions 1–2, CDT weakly dominates any decision rule whose implied act is constant across 12. Under Assumptions 1–2, PDT is globally optimal:
13
The proof sketch is pointwise: 14, and PDT chooses 15 maximizing 16 for each 17, so integration preserves the inequality (Sjölander, 29 Jun 2026).
These theorems delimit, rather than eliminate, controversy. They establish PDT’s optimality only under specific assumptions about subjective accuracy and about the absence of a direct pathway from decision theory to utility. The framework therefore clarifies which disputes are substantive disagreements about causal structure and which are merely terminological disagreements about what counts as the relevant expectation.
5. Canonical paradoxes, mechanised taxonomy, and controversy
The smoking–lesion problem is used to display the divergence between evidential, causal, and personal criteria. In Sjölander’s NPSEM, 18 denotes genetic lesion presence, 19 smoking or not, 20 cancer or not, and 21 utility, with equations
22
Here EDT chooses 23, CDT chooses 24, and PDT chooses 25. In many versions of the problem, 26 but 27, so EDT does not smoke, CDT does smoke, and PDT tailors the act to the particular 28 (Sjölander, 29 Jun 2026).
Mechanised causal graphs provide a broader taxonomy. Two axes generate six procedures: method—Evidential, Causal, Functional—and updatefulness—Updateful or Updateless. Standard EDT is updateful evidential choice,
29
Standard CDT is updateful causal choice,
30
Standard FDT is updateless functional choice,
31
The mechanised formalism emphasizes that the theories differ not only in what they optimize, but also in whether they condition on the realized action, intervene on the action, or intervene on the mechanism generating the action (MacDermott et al., 2023).
Newcomb’s problem exposes the importance of direct causal effects of the decision theory. In Sjölander’s NPSEM, variables include agent characteristics 32, decision theory 33, model 34, predictor’s prediction 35, act 36, and payoff 37. Under Assumption 1, and provided the predictor is sufficiently accurate so that 38, the analysis yields
39
Yet the same discussion notes that 40 gives a direct effect of 41 on 42, so the global optimality theorem for PDT need not apply (Sjölander, 29 Jun 2026). In the mechanised-graph formulation, by contrast, FDT is represented as an intervention on 43 in a logical-causal model, which reproduces one-boxing in Newcomb and cooperation in Twin Prisoner’s Dilemma when predictors or twins depend on the same mechanism (MacDermott et al., 2023). The resulting controversy is not erased by the formalism; rather, the formalism makes explicit which arrows, observations, and intervention targets produce each verdict.
6. Reflective, predictive, and adjacent extensions
A more human-centered line of work connects personal decision theory to metacognition rather than only to counterfactual choice rules. In this usage, personal decision theory emphasizes that people rarely make choices in a purely rational, utility-maximizing way; bounded cognitive resources, heuristics, and affective biases steer decision making toward solution-driven shortcuts. “Pre-Decision Reflection” shifts assistance from “here’s what you should do” to “here’s how you are thinking about it.” PROBE operationalizes this with two orthogonal metrics over seven reflective aspects—Belief, Awareness of Difficulties, Experience, Feeling, Intention, Insight, and Alternative Perspective. Breadth is
44
and depth is
45
In the reported sample, breadth ranged from 46 to 47 with mean 48 and 49, while 80% of participants had 50. Reliability was established with Fleiss’s 51 on a triple-coded subset and Cohen’s 52 on a double-coded subset (Tarvirdians et al., 5 Oct 2025).
Another extension models individualized choice probabilistically. In the Quantum Decision Theory account, the probability of choosing prospect 53 is
54
where 55 is the utility factor and 56 is the attraction factor. For binary choice between a gamble 57 and a sure option 58, the classical component is
59
while the attraction component takes the form
60
The interference phase is parameterized through framing, time pressure, memory, and need. In the reported experiments, the full QDT model achieved 61 on Dataset 1 and 62 on Dataset 2, outperforming the stated CPT and machine-learning baselines, with paired 63-test significance at 64 (Zhang et al., 2021).
A separate computational line combines symbolic utility discovery with language-based personalization. ATHENA first discovers group-level symbolic utility functions
65
for each alternative 66, and then initializes and optimizes an individual textual template
67
Final prediction is generated from the adapted template together with the utility-guided context. On Swissmetro, ATHENA achieved 68, 69, 70, and 71; on Vaccine, 72, 73, 74, and 75. The ablation study reported that removing symbolic discovery or semantic adaptation drops accuracy by 10–15 points (Zhao et al., 4 Nov 2025).
Related individual-level models broaden the surrounding landscape. The justifiability model posits a true preference 76 together with a family 77 of justifiable total orders; for any menu 78, the justifiable candidates are
79
and choice is
80
This captures decision making constrained by morality, rationality, or other virtues (Ridout, 2020). Hope-and-prepare preferences instead define an incomplete ordering under uncertainty in which 81 only when both a pessimistic minimum over one set of priors 82 and an optimistic maximum over another set of priors 83 favor 84 over 85:
86
and
87
This model is described as “hopes for the best” while “prepares for the worst” (Bardier et al., 2024).
Taken together, these works suggest a common research program: the relevant object is not merely expected utility averaged over a population, but decision quality indexed to the particular agent, that agent’s latent counterfactuals, subjective model, justificatory constraints, reflective profile, or semantic context. The main unresolved disputes then concern which causal graph is appropriate, what information may legitimately condition the choice rule, whether observational regularities should refine individualized counterfactuals, and how far individual decision support should optimize outcomes versus cultivate self-awareness.