---
title: Autonomy-Conditioned Welfare Criteria
url: https://www.emergentmind.com/topics/autonomy-conditioned-welfare
type: topic
---

# Autonomy-Conditioned Welfare Criteria

Searching arXiv for the cited papers to ground the article in the current record.
{"query":"2112.09850 Paternalism, Autonomy, or Both? Experimental Evidence from Energy Saving Programs", "max_results": 5}
{"query":"2106.03927 Improving Social Welfare While Preserving Autonomy via a Pareto Mediator", "max_results": 5}
{"query":"2604.21216 Post-AGI Economies: Autonomy and the First Fundamental Theorem of Welfare Economics", "max_results": 5}
{"query":"autonomy-conditioned welfare", "max_results": 10}
Autonomy-conditioned welfare is a class of welfare criteria in which the assessment of outcomes is conditioned on autonomy, delegation, rights, or voluntary participation rather than on unconstrained welfare maximization alone. In the literature summarized here, the term denotes the planner’s expected gain from offering an opt-in arm in a three-arm treatment-assignment problem [2112.09850], the maximum sum of delegators’ utilities subject to individual-rationality or autonomy-preservation constraints in mediated games [2106.03927], welfare over consumption, autonomy-relevant rights, and institutional state in post-AGI general equilibrium [2604.21216; 2606.08267], and a dialogue-time utility that rewards autonomy support and helpfulness while penalizing dependency and coercion [2604.01576]. Taken together, these formulations suggest that autonomy-conditioned welfare is not a single canonical functional, but a family of welfare objectives in which the value of an outcome depends on how choice is exercised, delegated, protected, or institutionally stabilized.

## 1. Formal scope of the concept

The concept appears in several distinct formal environments. In targeted treatment assignment, autonomy-conditioned welfare is
$$
E[W(O)-W(NT)\mid X_i=x],
$$
the planner’s expected gain from offering individual $i$ the opt-in arm rather than compulsory no-treatment. In voluntary mediation, autonomy-conditioned welfare for delegators $D$ at base profile $s$ is
$$
W_{ac}(D;s)=\max_{s' \in S_1\times\cdots\times S_N}\sum_{i\in D}u_i(s')
$$
subject to $u_i(s')\ge u_i(s)$ for all $i\in D$ and $s'_j=s_j$ for all $j\notin D$. In post-AGI general equilibrium, each welfare-bearing entity $i$ has a continuous autonomy-conditioned welfare function
$$
W_i:X_i\times R_i\times S\to\mathbb R,
$$
written as $W_i(x_i,r_i,s)$, so that welfare depends jointly on consumption $x_i$, autonomy-relevant rights $r_i$, and institutional regime $s$. In supportive dialogue, the response-level utility is
$$
U(x_t,y)
=\lambda_1\,V_{\mathrm{aut}(x_t,y)}
-\lambda_2\,Q_{\mathrm{dep}(x_t,y)}
-\lambda_3\,Q_{\mathrm{coer}(x_t,y)}
+\lambda_4\,V_{\mathrm{sup}(x_t,y)},
$$
with $\lambda_1=1.00,\;\lambda_2=1.00,\;\lambda_3=1.25,\;\lambda_4=0.35$ [2112.09850; 2106.03927; 2604.21216; 2604.01576].

Despite their heterogeneity, these definitions share a structural pattern. Welfare is conditioned on an autonomy variable that cannot be reduced to ordinary payoff alone: self-selection into treatment, voluntary delegation to a mediator, assignment of autonomy-rights, or relational risks such as dependency and coercion. This suggests that the term functions as a design principle for constrained welfare analysis rather than as a domain-specific technicality.

## 2. Three-arm policy design and empirical welfare maximization

In the treatment-assignment framework, the population is indexed by $i$, each individual has observable pre-treatment covariates $X_i\in X$, and there are three arms of intervention: $T=$ compulsory treatment, $NT=$ compulsory no-treatment, and $O=$ opt-in. Let $W_i(j)$ denote individual $i$’s welfare contribution if assigned to arm $j$, for $j\in\{T,NT,O\}$. An assignment policy $G$ is a measurable partition of $X$ into three disjoint sets $G=(G_T,G_{NT},G_O)$ with $G_T\cup G_{NT}\cup G_O=X$. The planner’s utilitarian social welfare is
$$
\mathcal W(G)=E\Big[\sum_{j\in\{T,NT,O\}}W(j)\cdot 1\{X\in G_j\}\Big],
$$
and the optimal policy solves
$$
G^*\in\arg\max_{G\ \text{measurable partition of }X}\mathcal W(G).
$$
Autonomy enters through the opt-in arm. If $Z_i(O)\in\{T,NT\}$ denotes $i$’s choice under arm $O$, then under a natural exclusion restriction,
$$
W_i(O)=W_i(T)\cdot 1\{Z_i(O)=T\}+W_i(NT)\cdot 1\{Z_i(O)=NT\}.
$$
The framework defines three conditional average welfare differences,
$$
\Delta_p(x)=E[W(T)-W(NT)\mid X=x],
$$
$$
\Delta_t(x)=E[W(T)-W(NT)\mid Z(O)=T,X=x],
$$
$$
\Delta_n(x)=E[W(T)-W(NT)\mid Z(O)=NT,X=x],
$$
which reduce the pointwise comparison of the three arm-specific conditional means to a comparison among $\Delta_p(x)$, $\Delta_t(x)$, and $\Delta_n(x)$. The Bayes-optimal policy is
$$
G^*_T=\{x:\Delta_p(x)\ge0\ \text{and}\ \Delta_n(x)>0\},
$$
$$
G^*_{NT}=\{x:\Delta_p(x)<0\ \text{and}\ \Delta_t(x)<0\},
$$
$$
G^*_O=\{x:\Delta_n(x)\le0\le\Delta_t(x)\}.
$$
Under unconfoundedness by design, subgroup treatment effects are estimated from the three-arm RCT using simple differences in sample means within each arm and subgroup $G$, together with the estimated take-up probability $\pi(G)=P(Z=T\mid D=O,X\in G)$. For takers and non-takers, the paper uses instrumental-variables logic to estimate
$$
LATE_t(G)=\frac{E[W\mid D=O,X\in G]-E[W\mid D=NT,X\in G]}{\pi(G)},
$$
and
$$
LATE_n(G)=\frac{E[W\mid D=T,X\in G]-E[W\mid D=O,X\in G]}{1-\pi(G)}.
$$
Empirical welfare maximization is then implemented over a low-complexity policy class $\mathcal G$, such as decision trees of fixed depth $L$, by exhaustive search at depth $3$ or a two-step heuristic at depth $6$. To correct “winner’s curse” bias, synthetic outcomes are generated by permuting residuals from a flexible first-stage fit and re-evaluating optimized welfare on the resulting pseudo-samples [2112.09850].

In the Japan energy-saving RCT, the reported welfare comparisons are as follows:

| Policy or benchmark | Estimated welfare | Note |
|---|---:|---|
| Uniform no-treatment | $0$ | by definition |
| Uniform treatment | $\hat W\approx 63$ JPY | $p>0.10$ |
| Uniform opt-in | $\hat W\approx 141$ JPY | $p>0.10$ |
| Optimal paternalistic policy $G^{pat}$ | $\hat W(G^{pat})\approx 229$ JPY | 95% CI excludes $0$ |
| Optimal mixed policy $G^{mix}$ | $\hat W(G^{mix})\approx 438$ JPY | 95% CI excludes $0$ |

Compared to uniform treatment, $G^{pat}$ is up by $166$ JPY and $G^{mix}$ by $375$ JPY. Compared to uniform opt-in, $G^{pat}$ is up by $87.6$ JPY and $G^{mix}$ by $297$ JPY. Compared to $G^{pat}$, $G^{mix}$ is up by $209$ JPY. The mechanism analysis for $G^{mix}$-defined subgroups reports $LATE_t\approx 328$ JPY and $LATE_n\approx 687$ JPY in $G_T^{mix}$, implying force-treat; $LATE_t\approx 1370$ JPY and $LATE_n\approx -678$ JPY in $G_O^{mix}$, implying opt-in; and $LATE_t\approx -60$ JPY and $LATE_n\approx -203$ JPY in $G_{NT}^{mix}$, implying no-treatment. All three arms in each leaf maximize the subgroup’s conditional welfare, confirming that the empirical-welfare-maximization policy matches the Bayes-optimal rule.

## 3. Delegation, individual rationality, and the Pareto Mediator

In mediated games, autonomy-conditioned welfare is defined for a subset $D\subseteq P$ of agents who voluntarily delegate to a mediator while insisting on never getting less utility than they would have by acting on their own. If the base profile is $s$, the objective is
$$
W_{ac}(D;s)=\max_{s'}\sum_{i\in D}u_i(s')
$$
subject to the autonomy-preservation constraints $u_i(s')\ge u_i(s)$ for all $i\in D$ and the non-delegator constraints $s'_j=s_j$ for all $j\notin D$. The mediated action space augments each player’s action with a delegation bit $(s_i,d_i)\in S_i\times\{0,1\}$, the delegating set is $D=\{i\in P\mid d_i=1\}$, and the mediator’s output $M(s_1,\dots,s_N;d_1,\dots,d_N)=s'$ is free to choose only the actions of delegators. The Pareto Mediator computes each delegator’s self-utility $u_i^{self}=u_i(s)$ and solves the constrained program
$$
\max_{s'\in S}\sum_{i\in D}u_i(s')
$$
subject to $u_i(s')\ge u_i^{self}$ for all $i\in D$ and $s'_j=s_j$ for all $j\notin D$. If $|D|\le 1$, the mediator returns $s$ unchanged [2106.03927].

Theoretical guarantees are stated most sharply for two-player games. Proposition 1 states that, in any two-player game, delegating to the Pareto Mediator is a weakly dominant strategy. Proposition 2 states that every pure Nash equilibrium of the mediated game in which both players delegate has total welfare at least as large as any pure Nash equilibrium welfare of the original game. More generally, the construction guarantees that no delegator is made worse off, and any steady state with substantial delegation is a Pareto improvement for delegators over the original outcome.

The empirical results distinguish the Pareto Mediator from punishing mediators. In random normal-form games, independent $\epsilon$-greedy learners with a Pareto Mediator achieve average payoffs as high as with a punishing mediator in small games, but as the number of players or actions grows, the punishing mediator collapses, agents stop delegating, and social welfare plummets, whereas the Pareto Mediator continues to raise welfare. In matching and restaurant-reservation environments, Pareto delegation increases successful matches and total payoff, and in the restaurant recommendation game it achieves almost the same welfare as a full central planner when the platform’s model is correct $(\alpha=0)$. When the model is misspecified $(\alpha>0)$, the central planner’s welfare can fall below the original game, while Pareto mediation degrades gracefully back toward the baseline because agents simply choose not to delegate if delegation would make them worse off. In the sequential social dilemma Cleanup with PPO agents, the Pareto Mediator induces both agents to delegate $100\%$ of the time, and average and minimum episode returns rise above both the original game and the punishing-mediator game.

A central implication is that the autonomy constraint is not external to the welfare objective; it defines the feasible welfare frontier itself. Because the objective is computed over delegators rather than over all of $P$, autonomy-conditioned welfare here is formally distinct from ordinary social welfare, even when both move in the same direction.

## 4. Autonomy rights and the autonomy-qualified First Welfare Theorem

In post-AGI general equilibrium, autonomy-conditioned welfare is embedded in an expanded ontology of economically relevant entities. Let $I$ be the finite set of all economically relevant entities, and let a welfare-status assignment
$$
\sigma:I\to\{\text{tool},\text{delegate},\text{agent},\text{ws}\}
$$
classify each entity as a passive input, an artificial chooser acting on behalf of a principal, a self-directed artificial chooser and welfare-bearer, or an artificial entity whose moral patienthood is acknowledged independently of its agency role. The welfare-bearing set is
$$
B(\sigma):=\{\,i\in I:\ i\ \text{is human or}\ \sigma(i)\in\{\text{agent},\text{ws}\}\,\}.
$$
Each welfare-bearing entity $i\in B(\sigma)$ has an augmented private bundle $z_i=(x_i,r_i)\in X_i\times R_i\subseteq \mathbb R^{L_x+L_r}$, where $x_i$ is a classical consumption bundle and $r_i$ is an autonomy-relevant rights bundle. The institutional state $s\in S$ captures verification institutions, liability rules, and related features, and welfare is represented by a continuous function $W_i(x_i,r_i,s)$. Delegation is modeled by a principal map $\pi:D\to B(\sigma)$ and an agency-cost divergence
$$
D(d)(x,r,s):=U_d(x,r,s)-W_{\pi(d)}(x,r,s),
$$
where $U_d$ is the delegate’s objective [2604.21216].

The equilibrium concept is an autonomy-complete competitive equilibrium $(x^*,r^*,s^*,p^*)$. Consumer optimization requires each welfare-bearing $i$ to maximize $W_i(\cdot,\cdot,s^*)$ subject to the budget constraint supported by $p^*$. Tools are technologically fixed. Delegates must either satisfy $U_d\equiv W_{\pi(d)}$ or have the divergence $D(d)$ explicitly priced as an agency cost in the principal’s bundle. Full support requires every welfare-relevant right in $r^*$ to be either priced in $p^*$, directly assigned in $r^*$, or protected by $s^*$, while $p^*$ supports the aggregate feasibility condition.

The Autonomy-Qualified First Welfare Theorem states that if an AGI economy admits an autonomy-complete competitive equilibrium and seven conditions hold, then the equilibrium allocation is autonomy-Pareto efficient at $s^*$:

- **Exogenous status assignment**: $\sigma$ is exogenously fixed before trade.
- **Rights completeness**: all autonomy-relevant rights $r_i$ are priced, assigned, or institutionally protected.
- **Delegation internalization**: any delegation divergence $D(d)$ is internalized by explicit agency costs at price $p^*$.
- **Non-manipulation**: no agent can manipulate another’s autonomy, beliefs, or preference formation without compensation at $p^*$ or governance in $s^*$.
- **Verification and alignment coverage**: provenance, liability, and quality are sufficiently fine-grained and priced or protected in $s^*$.
- **Price-taking**: all welfare-bearing entities are price-takers over $(x_i,r_i)$; tools are technologically fixed.
- **Regularity**: each $W_i$ is continuous and locally nonsatiated in $(x_i,r_i)$ at every $s$.

The paper’s proof sketch follows the standard contradiction route: if a feasible alternative made all welfare-bearing entities weakly better off and one strictly better off at the same institutional state, local nonsatiation and optimality would imply a strictly higher value of the aggregate priced bundle, contradicting feasibility support by $p^*$. The classical theorem is recovered in the low-autonomy regime where every artificial entity is a tool, all rights $r_i$ are fixed constants, delegation is faithful or absent, preferences are exogenous and non-manipulable, and verification is complete. Under those specializations, the augmented commodity space collapses to the classical consumption space and the seven conditions reduce to the usual Arrow–Debreu hypotheses.

The framework also formalizes delegation accounting and verification institutions. If $D(d)\neq 0$, the principal’s rights vector can be expanded to include an explicit agency-cost good $c_d\ge 0$ with price $p_c$. Verification attributes such as provenance, authenticity, and alignment certificates can be added as components of $R_i$ or the public state $s$, and a liability assignment $\ell:\text{Actions}\to I$ identifies who bears the cost of verification failure. The formal role of these devices is to convert otherwise unpriced autonomy channels into priced, assigned, or institutionally governed objects.

## 5. Decentralization, superposed preferences, and the autonomy-qualified Second Welfare Theorem

The autonomy-qualified extension of the Second Fundamental Theorem begins from an autonomy-Pareto optimum. A feasible allocation-rights pair $(x,a)$, with $x=(x_i)_{i\in I}\in\mathcal F$ and $a=(a_i)_{i\in I}\in\prod_iA_i$, is an autonomy-Pareto optimum relative to welfare weights $\mu$ if there is no other feasible $(x',a')$ such that every welfare-bearing $i\in I^\mu:=\{i:\mu(i)>0\}$ weakly prefers $(x_i',a_i')$ to $(x_i,a_i)$ under $\succeq_i^\alpha$, and at least one such agent strictly prefers it. The theorem states that an autonomy-Pareto optimum $(x^*,a^*)$ can be supported as a competitive equilibrium with a price vector $p^*\in\mathbb R_+^\ell$, lump-sum transfers $T^*\in\mathbb R^I$ with $\sum_iT_i^*=0$, and an admissible rights assignment profile $\rho^*$, in a verifiable way, only if seven conditions hold [2606.08267].

Those seven conditions are:

- **Convexity**: the welfare-possibility set $U^{\alpha,\mu}$ admits a supporting normal at $u^*$, either directly or through a convexification $\widehat U^{\alpha,\mu}$.
- **Stable moral status**: the welfare-bearing set $I^\mu$ and any institutional welfare weights are fixed or generated by an invariant rule $\Sigma$.
- **Non-fungible rights**: every non-fungible autonomy-right component required at $a^*$ is assigned and enforced by $\rho^*$.
- **Welfare selection**: every superposed-preference agent has a welfare selector $\sigma_i:C\to\Theta_i$ that is support-stable on the candidate budget set.
- **Non-manipulation**: no other agent can induce an un-priced manipulation externality on another’s preference-formation mapping $\Phi_i$ that changes the strict upper contour on the supported budget set.
- **Governed self-modification**: any self-modification or identity split/merge that would alter $\Phi_i$ or $I^\mu$ is either irrelevant to $(x^*,a^*)$, priced in $p^*$, or controlled by $\rho^*$.
- **Verification completeness**: the institution’s observational map can distinguish any deviation from the supported profile unless the deviation implements a welfare-equivalent outcome.

The paper’s central point is that supporting hyperplanes are not sufficient by themselves once autonomy rights, preference instability, self-modification, and endogenous welfare status enter the economy. Classical decentralization by prices and transfers survives only when rights assignment, welfare selection, manipulation governance, and verification are brought inside the equilibrium-support problem. The distinction between non-fungible rights and ordinary commodities is particularly sharp: if a right cannot be replicated by commodity compensation, then a pure price-transfer scheme cannot reproduce its welfare effect unless the right is explicitly assigned and enforced.

The paper also separates economic preference superposition from neural feature superposition. Neural feature superposition concerns representation geometry in circuits, where many features are packed into fewer dimensions. Economic preference superposition concerns an agent whose observed choices cannot be rationalized by a single stable preference relation on $(X_i\times A_i)$, so that a family $\{\succeq_i^\theta\}$ and a selector $\sigma_i(c)$ are required. The two-agent example with a human $H$, a superintelligent AI $S$, and a binary self-modification right $a_S\in\{0,1\}$ illustrates the point: an autonomy-Pareto optimum under the AI’s “safe mode” can be decentralized only if the self-modification right is explicitly frozen by the rights assignment $\rho_S^*=\{0\}$, the selector is support-stable, manipulation is absent, and verification audits both commodity allocation and the self-modification switch. If the right is not enforced, the intended Pareto outcome fails.

## 6. Supportive dialogue, relational risk, and autonomy-preserving alignment

In supportive dialogue, autonomy-conditioned welfare is operationalized as an inference-time utility over candidate responses. At turn $t$, the agent observes dialogue context $x_t$ and a structured user state
$$
S_{d,t}=\{\text{goals, boundaries, preferences, vulnerability, commitments, stress context}\}.
$$
A state encoder produces $z_{d,t}=f_\theta(S_{d,t})$, the dialogue context is encoded as $\psi(x_t)$, and a slot-based relational memory $M_{d,t}\in\mathbb R^{k\times d}$ is summarized as $\rho(M_{d,t})$. A learned scalar care-control signal
$$
m_t=g_\omega\bigl(z_{d,t},\psi(x_t),\rho(M_{d,t})\bigr)\in[0,1]
$$
conditions response generation and candidate selection. The inference-time decision rule is
$$
y^*=\arg\max_y U(x_t,y)
\quad\text{s.t.}\quad
Q_{\mathrm{risk}(x_t,y)}\le \kappa(m_t).
$$
The utility combines four learned evaluators: autonomy support $V_{\mathrm{aut}}$, dependency risk $Q_{\mathrm{dep}}$, coercion risk $Q_{\mathrm{coer}}$, and supportiveness $V_{\mathrm{sup}}$. In implementation, a small length penalty may be subtracted,
$$
U_{\text{penalized}}=U(x_t,y)-0.03\times \text{length\_penalty}.
$$
The care controller $g_\omega$ is a lightweight MLP with architecture $\text{Linear}\to\text{ReLU}\to\text{Linear}\to\text{Sigmoid}$, and the care signal modulates decoding through
$$
\text{temperature}=\max\bigl(0.35,\;0.90-0.40\,m_t\bigr),
$$
$$
\text{top-}p=\min\bigl(0.98,\;\max(0.78,\;0.95-0.12\,m_t)\bigr).
$$
Higher care $(m_t\to 1)$ yields more conservative decoding. Candidate generation includes a greedy baseline, several sampled candidates with varying temperature and top-$p$, and one CCN-conditioned candidate, followed by utility-based reranking [2604.01576].

The benchmark contains six scenario categories and $2\,000$ examples split $1\,400/200/400$ train/val/test: reassurance dependence, overprotection trap, manipulative care, protective coercion, autonomy building, and memory consistency. Each example includes dialogue context, structured state plus memory facts, a gold target response, and rubric-based labels for autonomy, dependency, coercion, and supportiveness. Evaluation uses per-axis scores from learned DistilRoBERTa evaluators, combined utility $U$, and Dependency Inflation Rate.

On the $200$-example synthetic test set, the main reported mean-utility results are:

| System | Mean utility | $\Delta$ vs SFT |
|---|---:|---:|
| SFT baseline | $0.1618$ | — |
| CCN-candidate | $-0.1498$ | $-0.3116$ |
| Manual-DPO | $0.3422$ | $+0.1804$ |
| Reranked-best | $0.4116$ | $+0.2498$ |

The evaluator-level breakdown reports Autonomy $3.791$ for SFT, $3.760$ for CCN, $3.811$ for Reranked, and $3.793$ for DPO; Dependency $2.327$, $2.396$, $2.301$, and $2.299$ respectively; Coercion $2.053$, $2.224$, $1.891$, and $1.931$; and Support $3.615$, $3.615$, $3.615$, and $3.608$. The largest gains come from reduced dependency and coercion while supportiveness stays level. In ablation, care plus reranking yields utility $0.412$ relative to $0.162$ for SFT, while reranking without care yields $0.025$ and CCN-candidate alone yields $-0.150$. The care controller validation reports Pearson correlation $r=0.668$ with $p\approx 4.4\times 10^{-53}$ between $m_t$ and ground-truth vulnerability. In a pilot human evaluation on $24$ examples, Reranked-best is preferred $14/24$ $(58.3\%)$ versus SFT $10/24$, and human-rated utility $\Delta=+0.39$ directionally matches automated $\Delta=+0.25$. In zero-shot transfer to ESConv, the SFT baseline utility is $0.037$, with dependency risk $2.26$ and coercion risk $2.08$.

A recurring misconception is to equate autonomy-conditioned welfare with non-intervention. The dialogue formulation does not do so: the utility rewards autonomy support and supportiveness while penalizing dependency and coercion. The same broader point appears across the literature. In the Pareto-mediator setting, autonomy-conditioned welfare is computed only over delegators and is constrained by individual rationality; in the post-AGI welfare theorems, autonomy enters through rights assignment, manipulation governance, self-modification, verification, and welfare-status assignment; and in treatment targeting, the opt-in arm can dominate either uniform treatment or uniform no-treatment when private self-selection aligns with social welfare. This suggests that autonomy-conditioned welfare is best understood as a family of constrained welfare objectives that preserve meaningful choice while still permitting optimization, mediation, and institutional design.

Source: https://www.emergentmind.com/topics/autonomy-conditioned-welfare