---
title: Subsidy-Sorting Principle
url: https://www.emergentmind.com/topics/subsidy-sorting-principle
type: topic
---

# Subsidy-Sorting Principle

The subsidy-sorting principle denotes a class of subsidy rules in which subsidy assignment is used to rank, segment, or induce self-selection among heterogeneous agents rather than to treat all agents identically. In the cited literature, the principle appears in several distinct but related forms: ranking ride-hailing queries by predicted uplift per subsidy dollar under a global budget constraint; choosing personalized continuous subsidies that induce treatment for those with positive marginal treatment effect; allocating free vaccination on a network where imitation can amplify or reverse the intended targeting; ordering sequential search by subsidies that signal product quality; and sorting generation investment by each producer’s marginal contribution to consumer surplus [2408.02065] [2202.13545] [1503.08048] [2605.28985] [2407.02161].

## 1. General statement and recurring structure

Across these formulations, the object being sorted differs, but the operative logic is consistent: subsidies are attached to heterogeneous units, and the resulting ordering determines who is induced, inspected, vaccinated, or invested in first. In some models the ordering is explicit, as in “ranking or segmenting consumers according to their estimated marginal benefit (uplift) per dollar of subsidy.” In others it is implemented indirectly through self-selection, as when a subsidy induces all individuals with sufficiently low resistance into treatment, or when higher-quality firms choose weakly larger subsidies and are therefore searched first [2408.02065] [2202.13545] [2605.28985].

| Setting | Sorted object | Sorting rule or outcome |
|---|---|---|
| Multi-class ride-hailing | queries or clusters | estimated marginal benefit (uplift) per dollar of subsidy |
| Personalized subsidy rules | individuals indexed by $(x,u)$ | individuals are induced in decreasing order of their MTE |
| Vaccination on networks | nodes under limited free vaccination | targeted subsidy depends on degree but interacts with imitation bias |
| Sequential search | firms | higher-quality firms provide weakly larger subsidies |
| Renewable generation investment | producers or technologies | sort-by-$\Delta \mathrm{CS}$ |

The principle is therefore not a single theorem with one universal mathematical representation. It is a family of subsidy designs in which a subsidy schedule creates an ordering over heterogeneous agents, and the ordering is then used to implement a welfare, revenue, epidemic-control, search, or investment objective.

## 2. Personalized subsidy rules and marginal-treatment-effect ordering

In Chen and Xie’s formulation, the primitive objects are a binary treatment $D \in \{0,1\}$, potential outcomes $Y_0,Y_1$, covariates $X$, and an unobserved resistance $U_D \sim \mathrm{Unif}[0,1]$ entering selection through
$$
D=1\{g(X,W,Z) \ge U_D\},
$$
where $Z$ is the continuous subsidy and $W$ may be auxiliary instruments. The Marginal Treatment Effect is
$$
MTE(x,u) \equiv E[Y_1-Y_0 \mid X=x,U_D=u].
$$
It is interpreted as the increment in average outcome from switching $D=0 \to 1$ among individuals with observables $x$ and indifference-position $u$; equivalently, it is the local average treatment effect at the margin $u$ [2202.13545].

A personalized subsidy rule $s(\cdot)$ maps each realized covariate vector $x$ to a subsidy amount $z=s(x)$. Under this rule,
$$
D^s=1\{g(X,W,s(X)) \ge U_D\},
$$
and the realized outcome is $Y^s=D^sY_1+(1-D^s)Y_0$. If $c(x,z,d)$ is the per-individual cost of giving subsidy $z$ to an $x$-type who then selects $d \in \{0,1\}$, the average cost is $C(s)=E[c(X,s(X),D^s)]$, and welfare is
$$
W(s)=E[Y^s]-E[C(s)].
$$
In the important case $c(x,z,d)=z \cdot d$, one obtains
$$
E[Y^s]=E[Y_0]+E\Bigl[\int_0^{g(X,s(X))} MTE(X,u)\,du\Bigr],
$$
$$
E[C(s)]=E[s(X)\cdot g(X,s(X))].
$$
Hence
$$
W(s)=E[Y_0]+E\Bigl[\int_0^{g(X,s(X))} MTE(X,u)\,du-s(X)\cdot g(X,s(X))\Bigr].
$$

Fixing $x$ and writing $p(s)=g(x,s)$, the inner objective is
$$
V_x(s)\equiv \int_0^{p(s)}MTE(x,u)\,du - s\cdot p(s).
$$
Differentiation yields the pointwise first-order condition
$$
MTE(x,p(s))\cdot p'(s)-[p(s)+s\cdot p'(s)] = 0,
$$
or equivalently,
$$
MTE(x,p(s)) = s + \frac{p(s)}{p'(s)}.
$$
The subsidy-sorting principle in this setting is that “Individuals are induced in decreasing order of their MTE.” Writing $u^*(x)\equiv g(x,s^*(x))$, the positive-selection case is defined by $u \mapsto MTE(x,u)$ being weakly decreasing. Then the chosen cutoff satisfies
$$
MTE(x,u^*(x))=0,
$$
so all those with $MTE(x,u)>0$ are in treatment and those with $MTE(x,u)<0$ are out.

This formulation yields a first-best result under positive selection. If the planner could observe $(x,u)$, the full-information rule would be
$$
D^*(x,u)=1\{MTE(x,u)\ge 0\},
$$
with welfare
$$
FB(x)=\int_0^1 \max\{MTE(x,u),0\}\,du.
$$
When $MTE(x,u)$ is weakly decreasing in $u$, a subsidy $s^*(x)$ that induces take-up $g(x,s^*)=u^*$ achieves
$$
W(s^*)=E[Y_0]+E\Bigl[\int_0^{u^*(X)}MTE(X,u)\,du\Bigr]
      =E[Y_0]+E\Bigl[\int_0^1 \max\{MTE(X,u),0\}\,du\Bigr].
$$
The paper also distinguishes point identification when the MTE is fully known, point identification under positive selection without large-support as long as the relevant crossing lies inside support, and partial identification under shape restrictions. In the Jordan New Opportunities for Women pilot study, the authors estimate $g(x,z)$ and $MTE(x,u)$ from a parametric selection model and report that medical majors receive a positive $s^*$ well below the experimental maximum, while others have corner $s^*=0$.

## 3. Multi-class ride-hailing: causal ranking and budget-aware assignment

In the ride-hailing system, each arriving query is treated as a candidate for one of several discrete subsidy treatments. Let $X$ denote the feature vector for a query, $T \in \{0,1,\dots,J\}$ the discrete subsidy level, and $Y$ the observed outcome, where order $=1$ or $0$. The generalized propensity score is
$$
e_j(X)=P[T=j \mid X], \qquad j=0,\dots,J,
$$
the no-subsidy response is
$$
\mu_0(X)=E[Y(0)\mid X],
$$
and the uplift is
$$
\tau_j(X)=E[Y(j)-Y(0)\mid X], \qquad j=1,\dots,J.
$$
Predicted conversion under treatment $j$ is
$$
\mu_j(X)=\mu_0(X)+\mathrm{softplus}(\tau_j(X)).
$$
The system trains a network to output $\hat e_j(X)$, $\hat \mu_0(X)$, and $\hat \tau_j(X)$, and then uses them to compute each query’s predicted uplift per subsidy dollar [2408.02065].

MulTeNet consists of three modules branching from a shared “feature net” $h=f_\phi(X)$: a GPS Net $g_\psi(h)\to \mathrm{softmax}\to \{\hat e_j\}$ estimating $P[T=j\mid X]$; a $Y_0$ Net $m_\theta(h)\to \hat \mu_0(X)$; and a Monotone Net $u_\omega(h)\to \{\hat \tau_j(X)\}$ with softplus activations to enforce non-negative, monotonic uplift. The training loss is
$$
L=L_{\mathrm{outcome}}+\alpha L_{\mathrm{prop}}+\lambda L_{\mathrm{orth}},
$$
with
$$
L_{\mathrm{outcome}}
=\frac{1}{N}\sum_i\Bigl[Y_i-\bigl(\hat \mu_0(X_i)+\mathrm{softplus}(\hat \tau_{T_i}(X_i))\bigr)\Bigr]^2,
$$
$$
L_{\mathrm{prop}}
=-\frac{1}{N}\sum_i \log \hat e_{T_i}(X_i),
$$
and
$$
L_{\mathrm{orth}}
=\text{orthogonal regularizer }(h,\nabla_h \log e,\nabla_h(\mu_0+\tau)).
$$
The orthogonal term, following Hatt and Feuerriegel, penalizes correlations between the gradient of the propensity model and that of the outcome model in the shared representation $h$, thereby reducing confounding bias.

The allocation pipeline is explicitly two-stage. Offline, every few hours or daily, MulTeNet is trained on all historical queries $\{X_i,T_i,Y_i\}$; the feature space is clustered, for example by origin, destination, and time; and for each cluster $i$ and treatment $j$ the system computes
$$
\hat P_{i,j}=\mathrm{average}_{X \in \mathrm{cluster}\ i}\,[\mu_0(X)+\tau_j(X)],
$$
$\hat{Pr}_{k,i}$ as expected revenue if service class $k$ is chosen, and $\hat N_i$ as the forecasted number of queries in cluster $i$. It then solves the budget-constrained optimization
$$
\max_{x_{i,j}\in\{0,1\}}
\sum_{k,i,j}\gamma_k\,\hat N_i\,\hat{Pr}_{k,i}\,\hat P_{i,j}\,x_{i,j}
$$
subject to
$$
\sum_{k,i,j}\gamma_k\,\hat N_i\,\hat P_{i,j}\,x_{i,j}\,c_{k,i,j}\le B,
$$
$$
u_{\mathrm{lower}} \le \sum_j \hat P_{i,j}\,x_{i,j}\,c_{k,i,j}\le u_{\mathrm{upper}}, \qquad \forall\,i,k,
$$
$$
\sum_j x_{i,j}=1, \qquad \forall\,i.
$$
The output is a lookup table that records, for each cluster $i$ and service class $k$, the optimal subsidy index $j^*$. Online, the system extracts features $X$, identifies cluster $i$ and service class $k$, looks up $j^*$, and presents the user the subsidy level $j^*$. Because the heavy causal inference and optimization are done offline, the online latency is just a hash lookup.

The causal assumptions are conditional ignorability, overlap, and SUTVA / no interference across queries. Bias mitigation is handled by explicitly modeling and penalizing the generalized propensity scores through $L_{\mathrm{prop}}$ and by applying orthogonal regularization through $L_{\mathrm{orth}}$. Offline evaluation on $\sim 2\,\mathrm{M}$ queries, with $200\,\mathrm{K}$ subsidized, used AUC, AUUC, and the QINI coefficient against transfer-learning (Yu et al 2023) and DragonNet. MulTeNet achieved the best AUUC, $1.936$ versus $1.809/1.382$, and the best QINI, $0.325$ versus $0.216/0.202$. In a $7$-day online budget-constrained experiment at a $5\%$ target subsidy rate, the reported outcomes under $\approx 5\%$ subsidy were revenue up $5.47\%$, orders up $2.98\%$, and ROI $=1.05$, with the text stating “+34% over best baseline.”

## 4. Vaccination on complex networks: node targeting, imitation, and reversal effects

In the vaccination model, a fixed budget allows exactly a fraction $p$ of the population to receive a free vaccine, while the remainder decide voluntarily whether to vaccinate. Targeted subsidy chooses the $pN$ nodes of highest degree and gives them vaccine at zero personal cost. Random subsidy chooses $pN$ nodes uniformly at random. For non-subsidized individuals, vaccination costs $C_V$. The epidemic then follows an SIR process with per-contact transmission rate $\lambda$ and recovery rate $\mu$, and the relative vaccination cost is
$$
c=\frac{C_V}{C_I}\in[0,1],
$$
with $C_I=1$ as the baseline unit [1503.08048].

The behavioral layer is an imitation rule. Each non-subsidized node chooses a neighbor as an “imitation object” with probability
$$
\Pr\{i\to j\}
=\frac{\exp(\alpha w_j)}
{\sum_{k\in \mathcal N(i)}\exp(\alpha w_k)},
$$
where
$$
w_j=
\begin{cases}
5,& j\in \mathcal S,\\
1,& j\notin \mathcal S.
\end{cases}
$$
If node $i$ compares with node $j$, it adopts $j$’s vaccination choice with probability
$$
W(s_i\leftarrow s_j)=\frac{1}{1+\exp[-\beta(P_j-P_i)]},
$$
with $\beta=10$. The sign of $\alpha$ determines whether non-subsidized nodes preferentially look at subsidized neighbors, non-subsidized neighbors, or select neighbors randomly.

The analytic representation combines a degree-based mean-field for vaccination with bond-percolation for the epidemic. Let $q_k(p,\alpha)$ be the probability that a degree-$k$ node is immune before the epidemic, either because it was pre-subsidized or because it vaccinated voluntarily:
$$
q_k(p,\alpha)=p_k+(1-p_k)V_k(p,\alpha).
$$
With transmissibility
$$
T=1-e^{-\lambda/\mu},
$$
and $u$ denoting the probability that a random edge does not transmit infection to the node at its far end, the locally tree-like approximation gives
$$
u = 1 - T
+ T \sum_{k}\frac{kP(k)}{\langle k\rangle}
\bigl[1-q_k(p,\alpha)\bigr]
u^{k-1}.
$$
The final epidemic size is then
$$
R_\infty(p,\alpha)
=\sum_k P(k)\bigl[1-q_k(p,\alpha)\bigr]\bigl[1-u^k\bigr].
$$

The central result is that degree-based subsidy sorting is not universally advantageous once imitation is endogenous. For $\alpha>0$, targeted subsidy outperforms random subsidy:
$$
R_\infty^T(p,\alpha)<R_\infty^R(p,\alpha).
$$
For $\alpha<0$, the ranking reverses:
$$
R_\infty^T(p,\alpha)>R_\infty^R(p,\alpha).
$$
The abstract states that the targeted strategy is only advantageous when individuals prefer to imitate the subsidized individuals’ strategy; otherwise, its effect is worse than random immunization. More strongly, under the targeted subsidy policy, increasing the proportion of subsidized individuals may increase the final epidemic size.

The welfare object is social cost,
$$
S(p,\alpha)
= c\Bigl[\sum_k P(k)q_k(p,\alpha)\Bigr] + R_\infty(p,\alpha),
$$
or equivalently
$$
S(p,\alpha)
= c\sum_k P(k)q_k(p,\alpha)
+\sum_k P(k)\bigl[1-q_k(p,\alpha)\bigr]\bigl[1-u^k\bigr].
$$
The paper reports that there exist some optimal intermediate regions leading to the minimal social cost. In the fuller exposition, the worst region is said to lie typically around $\alpha\in[-3,-1]$ and $p\in[0.03,0.1]$, where targeted subsidy can increase epidemic burden above even the no-subsidy case.

## 5. Sequential search: subsidy as quality signal and search-order rule

In the sequential-search model, there are $n\ge 1$ firms. Firm $j$ has privately known quality $q_j\in[0,1]$, often written $t$, drawn i.i.d. from a continuous distribution $F$ with density $f$. If the consumer inspects firm $j$, it matches with probability $t_j$; a successful match yields payoff $1$ to the firm. The consumer obtains gross utility $u>0$ from a match, each inspection carries gross cost $c>0$, and firms may choose subsidies $s\in[0,c]$. If the consumer inspects that firm, she pays net cost $c-s$, while the firm pays $p\cdot s$ out of its pocket [2605.28985].

The equilibrium concept is symmetric Perfect Bayesian Equilibrium refined by an equilibrium-dominance argument in the spirit of the Intuitive Criterion. A type-$t$ firm choosing subsidy $s$ earns
$$
\pi(t,s)=\bigl(t-p\,s\bigr)\cdot q(s),
$$
where $q(s)=\Pr\{\text{firm offering }s\text{ is inspected}\}$. The subsidy-sorting theorem states that in every symmetric PBE: $\sigma(t)$ is weakly increasing in $t$; the induced inspection probability $q(s)$ is weakly increasing in $s$; and the consumer inspects firms in order of highest subsidy first, breaking ties uniformly, and stops once further inspection yields negative net expected payoff.

Writing
$$
\tau(s)=E[t\mid \sigma(t)=s]
$$
for the posterior match probability, the reservation index is
$$
r(s)=u-\frac{c-s}{\tau(s)}.
$$
The consumer’s optimal search rule orders firms in decreasing $r(s)$, equivalently in decreasing $s$ when $\sigma$ is increasing. The result is explicitly described as descending-subsidy search.

Under the Intuitive-Criterion refinement, the equilibrium has three regions. Define
$$
\underline t = \frac{p\,c}{1+p}.
$$
For $t<\underline t$, firms choose $\sigma(t)=0$ and are never inspected. For $\underline t \le t \le \bar t$, the schedule is strictly increasing and satisfies
$$
p\,\frac{d}{dt}\bigl[\sigma(t)\,q^{\rm sep}(t;\underline t)\bigr]
= t\,\frac{d}{dt}q^{\rm sep}(t;\underline t),
$$
with boundary condition $\sigma(\underline t)=\underline t/p$, where
$$
q^{\rm sep}(t;\underline t)
=\Bigl(1-\int_t^1 x\,dF(x)\Bigr)^{n-1}.
$$
The closed form is
$$
\sigma^{\rm sep}(t)
=\frac{t}{p}
-\frac{1}{p\,q^{\rm sep}(t;\underline t)}
\int_{\underline t}^{t} q^{\rm sep}(x;\underline t)\,dx.
$$
For $t>\bar t$, all firms pool at the full subsidy $s=c$. This “step–increasing–step” equilibrium is stated to maximize information revelation among all PBE outcomes and to ensure efficient inspection.

The platform extension introduces inspection “tokens” sold at linear price $p$. Platform revenue is
$$
R(p)=n\cdot p\cdot D(p),
\qquad
D(p)=E_t[\sigma^*(t;p)\cdot q^*(t;p)].
$$
The optimal linear pricing has two reported features: pooling remains active, so $\bar t<1$, and some types with negative virtual value are inspected. The conclusion is that the platform’s optimal linear pricing leads to excessive inspection relative to the social optimum. The abstract adds that this distortion does not reduce consumer welfare, but reallocates surplus from sellers to the platform and consumers.

## 6. Renewable-generation investment: sorting by marginal contribution to consumer surplus

In the electricity-market setting, the subsidy-sorting idea is formulated as “sort-by-$\Delta \mathrm{CS}$.” For a dispatch interval $t$ and bus $n$, price $P_{tn}$ equals the marginal willingness-to-pay at dispatched quantity
$$
d_{tn}=D_{tn}(P_{tn})
=\arg\!\max_{0\le d\le D_{tn}^{\max}}
\{\tilde U_{tn}(d)-P_{tn}d\}.
$$
Consumer surplus is
$$
\mathrm{CS}_{tn}(P_{tn})
=\int_0^{D_{tn}(P_{tn})}\tilde U'_{tn}(x)\,dx
- P_{tn}D_{tn}(P_{tn}),
$$
and total consumer surplus is
$$
\mathrm{CS}(\{P_{tn}\})
=\sum_{t=1}^T\sum_{n\in\mathcal N}
\Bigl[
\int_0^{D_{tn}(P_{tn})}\tilde U'_{tn}(x)\,dx
- P_{tn}D_{tn}(P_{tn})
\Bigr].
$$
The subsidy is then tied to each producer’s marginal contribution to this surplus [2407.02161].

If producer $i$ delivers $q_{t\,i\,n}$ at $(t,n)$, its contribution at time $t$ is defined as
$$
\Delta \mathrm{CS}_{t,i}
=\sum_{n\in\mathcal N}
\Bigl[
U_t(q_{tn})-U_t(q_{tn}-q_{t\,i\,n})-P_{tn}q_{t\,i\,n}
\Bigr],
$$
and the full marginal contribution is
$$
\Delta \mathrm{CS}_i
=\sum_{t=1}^T \Delta \mathrm{CS}_{t,i}
=\sum_{t,n}
\bigl\{
U_t(q_{tn})-U_t(q_{tn}-q_{t\,i\,n})-P_{tn}q_{t\,i\,n}
\bigr\}.
$$
In the continuous-investment approximation, the subsidy rule is
$$
s_i=\int_0^{k_i}\frac{\partial}{\partial k_i}
\bigl[\mathrm{CS}(P(k))\bigr]\,dk_i,
\qquad
\frac{\partial s_i}{\partial k_i}
=\frac{\partial \mathrm{CS}}{\partial k_i}.
$$
In the discrete-output formulation, the payment is written as
$$
\chi_{ti}
=\sum_n
\Bigl[
U_t(q_{tn})-U_t(q_{tn}-q_{t\,i\,n})-P_{tn}q_{t\,i\,n}
\Bigr],
$$
with $\chi_{ti}=0$ when $q_{t\,i\,n}=0$.

The implementation claim is exact. Producer $i$ invests as long as its marginal investment cost satisfies
$$
\frac{\partial \mathfrak C_i}{\partial (\Delta k_i)}
\le \Delta \mathrm{CS}_i.
$$
If all technologies are listed in descending order of $\Delta \mathrm{CS}_i$ and allowed to invest until $\Delta \mathrm{CS}_i$ equals marginal investment cost, the result reproduces the planner’s first-order condition for socially optimal $k^*$. The exposition calls this a natural “subsidy auction” in which firms with the highest consumer-surplus-contribution win first.

A further feature of this formulation is informational. To compute $\chi_{ti}$, the regulator needs only the aggregate demand curve or consumer-utility function, the realized nodal price $P_{tn}$, and dispatch quantities $q_{t\,i\,n}$. The text states that no private cost-parameter vector of firm $i$ and no unobserved technology characteristics enter $\chi_{ti}$, so the regulator’s informational burden is identical to that of ordinary competitive clearing.

## 7. Assumptions, limits, and recurring controversies

The cited formulations do not treat subsidy-sorting as automatically welfare-improving under all environments. In the personalized-subsidy framework, first-best attainment depends on positive selection, meaning that $u\mapsto MTE(x,u)$ is weakly decreasing for each $x$; without that condition, the clean threshold characterization does not deliver the same conclusion [2202.13545]. In the ride-hailing system, unbiased uplift estimation depends on conditional ignorability, overlap, and SUTVA / no interference across queries, and the architecture adds generalized propensity modeling and orthogonal regularization precisely because confounding effects pose challenges in achieving an unbiased estimate of the uplift effect [2408.02065].

A common misconception is that targeting the apparently highest-risk or highest-value units must dominate untargeted allocation. The vaccination model directly rejects that claim: targeted subsidy is only advantageous when individuals prefer to imitate the subsidized individuals’ strategy, and under the opposite imitation bias it can perform worse than random immunization and may even increase final epidemic size as the subsidized fraction rises [1503.08048]. The principle is therefore sensitive to endogenous behavioral response, not merely to ex ante ranking.

Another recurrent issue concerns efficiency versus induced distortions. In sequential search, the refined equilibrium maximizes information revelation among PBE outcomes and ensures efficient inspection, yet the platform’s optimal linear pricing leads to excessive inspection relative to the social optimum [2605.28985]. By contrast, the renewable-investment scheme is presented as aligning private and social incentives without increasing the regulator’s information burden [2407.02161]. This suggests that the welfare properties of subsidy-sorting depend not only on the ranking criterion but also on who sets the subsidy schedule, what information the schedule reveals, and whether strategic or social-learning responses feed back into the allocation.

Taken together, these literatures define the subsidy-sorting principle as a structured use of subsidies to create an economically meaningful order: decreasing order of MTE, descending order of uplift per subsidy dollar, descending subsidy order in search, degree-based targeting under behavioral spillovers, or descending order of $\Delta \mathrm{CS}_i$. The substantive content of the principle lies in how that order is constructed, what assumptions justify it, and whether the induced ordering coincides with the relevant welfare or revenue objective in the environment under study.

Source: https://www.emergentmind.com/topics/subsidy-sorting-principle