---
title: Under Matching in Dynamic Markets
url: https://www.emergentmind.com/topics/under-matching
type: topic
---

# Under Matching in Dynamic Markets

*“Under-matching” (Editor’s term)* denotes a class of matching phenomena in which a mechanism does not execute every immediately feasible match. In the dynamic two-sided market studied in "Dynamic Matching Under Patience Imbalance" [2602.03995], the clearest instance arises when an \(L\)-type demand arrival could be matched with available \(H\)-type supply, yet the planner deliberately leaves that demand unmatched in order to preserve scarce high-quality supply for future \(H\)-type demand. Closely related ideas appear in stochastic matching with recourse, where an optimal policy may avoid a maximal matching in order to preserve future flexibility, and in sequential matching under uncertainty, where immediate assignments must be balanced against information acquisition about latent compatibility [1605.08616], [1707.09678].

## 1. Formal setting and basic mechanism

The canonical formalization of under-matching in the supplied literature is an infinite-horizon dynamic matching market with two sides, supply and demand, in discrete time. In each period, exactly one supply agent and one demand agent arrive. Each arrival is either high type \(H\) or low type \(L\); a newly arrived supply agent is \(H\) with probability \(p\) and \(L\) with probability \(1-p\), while a newly arrived demand agent is \(H\) with probability \(q\) and \(L\) with probability \(1-q\). Matching payoffs are denoted \(r_{ij}\) for supply type \(i\in\{H,L\}\) and demand type \(j\in\{H,L\}\), under the assumptions
\[
r_{HH}\ge r_{HL}\ge r_{LL}, \qquad r_{HH}\ge r_{LH}\ge r_{LL},
\]
together with supermodularity,
\[
r_{HH}+r_{LL}\ge r_{HL}+r_{LH}.
\]
The supply side is long-lived and incurs a per-period waiting cost \(h>0\), while the demand side is short-lived and departs if not matched immediately [2602.03995].

This patience asymmetry creates the possibility of deliberate nonmatching. Because demand does not queue, there is at most one match per period. A feasible match can therefore be postponed only in a specific sense: an arriving demand agent could be served by some available supply agent, but the planner or decentralized agents choose not to consummate the match. In the centralized model, the relevant post-arrival state is
\[
\mathbf{s}_t=(x_H,x_L,y_H,y_L),
\]
where \(x_H,x_L\in\mathbb Z_+\) count waiting \(H\)- and \(L\)-type supply and \(y_H,y_L\in\{0,1\}\) encode the current demand type. The action set is
\[
\mathbf{a}_t\in\{u_\phi,u_{HH},u_{HL},u_{LH},u_{LL}\},
\]
where \(u_\phi\) means no match. Under-matching, in this formulation, is not mere delay on both sides; it is the intentional sacrifice of current feasible trade because only the supply side can carry value into the future [2602.03995].

## 2. Efficient under-matching in centralized systems

In the centralized benchmark, the planner maximizes long-run average social welfare, defined as current matching payoff minus waiting costs for unmatched supply. The per-period reward is
\[
R(\mathbf{s}_t,\mathbf{a}_t)=\mathbf{r}\cdot \mathbf{a}_t - h(x_H+x_L-\mathbf{1}_{\mathbf{a}_t\neq u_\phi}),
\]
and the planner’s objective is
\[
W(\pi)\equiv \lim_{T\rightarrow \infty}\frac{1}{T}\mathbf{E}\left[\sum^{T}_{t=1}R(\mathbf{s}_{t}, \pi(\mathbf{s}_{t}))\right].
\]
The key structural result is a threshold policy. If an \(H\)-type demand arrives, the planner matches greedily: first with \(H\)-type supply if available, otherwise with \(L\)-type supply. If an \(L\)-type demand arrives, the planner first uses \(L\)-type supply if available; if no \(L\)-type supply is available, the planner uses \(H\)-type supply only when the number of waiting \(H\)-type suppliers exceeds a threshold \(k^{ce}\) [2602.03995].

This threshold is the formal expression of efficient under-matching. When \(x_H\le k^{ce}\), a feasible \((H,L)\) match is rejected and the \(L\)-type demand departs unmatched. The mechanism is entirely dynamic. Supermodularity creates an option value from preserving \(H\)-type supply for future \(H\)-type demand, while waiting cost \(h\) works in the opposite direction. The paper defines the supermodularity gap
\[
r:=r_{HH}+r_{LL}-r_{LH}-r_{HL},
\]
and shows that the optimal threshold \(k^{ce}\) increases in \(r\) and decreases in \(h\). This establishes a precise comparative-static description: stronger assortative gains imply more rationing of high-quality supply, while higher waiting cost weakens the incentive to leave feasible matches unrealized [2602.03995].

A central misconception is that under-matching necessarily reflects inefficiency or coordination failure. In this model, the opposite is true in the centralized benchmark: selective refusal to match can be first-best. The planner is not leaving value on the table in a static sense; it is exchanging current low-value trade for future high-value assortative trade.

## 3. Decentralized under-matching, private incentives, and misallocation

The decentralized version of the same market replaces planner control with bilateral acceptance and a fixed payoff split. Types are publicly observable, and a matched pair with total payoff \(r_{ij}\) allocates share \(\alpha\) to supply and \(1-\alpha\) to demand. Supply agents can wait and pay cost \(h\); demand agents are short-lived and receive zero if unmatched. The equilibrium concept is pure-strategy Markov perfect equilibrium [2602.03995].

Private waiting incentives generate a decentralized threshold
\[
k^{de}=\left\lfloor \frac{q\alpha (r_{HH}-r_{HL})}{h}\right\rfloor.
\]
The first \(k^{de}\) \(H\)-type supply agents prefer to refuse current \(L\)-type demand and wait for future \(H\)-type demand. There is also an \(L\)-type supply threshold \(k_L(x_H)\), which decreases in the current stock of \(H\)-type supply, satisfies \(k_L(x_H)\le k^{de}-x_H\), and equals \(0\) when \(p\ge q\). In the welfare-maximizing equilibrium, all supply is willing to match with \(H\)-type demand, but \(L\)-type demand is served only after these waiting incentives are respected [2602.03995].

The distinctive equilibrium phenomenon is that low-type demand may match with high-type supply even when low-type supply is available. This is not the centralized pattern. It occurs because some \(L\)-type supply agents strategically decline current \(L\)-type demand in order to wait for future \(H\)-type demand, while sufficiently deep \(H\)-type supply agents are willing to accept the current \(L\)-type demand. The resulting \((H,L)\) match is therefore a decentralized distortion: under-matching on one margin induces misallocation on another. The paper treats this as a defining difference between decentralized one-sided-backlog markets and both the centralized benchmark and full-backlog environments [2602.03995].

The same analysis also yields an implementation result. Because only supply accumulates, the payoff share \(\alpha\) can be used to tune private waiting incentives, and the decentralized system can be perfectly aligned with the centralized optimum by choosing \(\alpha\) appropriately. A plausible implication is that one-sided patience makes under-matching unusually controllable: the platform can regulate the only queue that matters.

## 4. Under-matching as preservation of future optionality

A different route to under-matching appears in "Maximum-expectation matching under recourse" [1605.08616]. There the environment is a stochastic matching problem on an undirected graph \(G=(V,E)\), motivated by kidney exchange restricted to 2-cycles. Edges may fail with probabilities \(p_{ij}\), and after observing successes and failures one may rematch on the residual graph for up to \(N\) recourse rounds. The objective is to maximize the expected number of matched vertices, computed recursively by evaluating each chosen matching under all success-failure patterns and solving again on the residual graph.

The central structural result is that, under recourse, a maximum-expectation matching may be non-maximal. The paper’s \(C_4\) example compares the maximal matching \(\{\{1,2\},\{3,4\}\}\) with the non-maximal matching \(\{\{1,2\}\}\), and shows that the latter can have higher expected value under recourse. The difference is
\[
-2 p_{12} (1 - p_{14}) (1 - p_{23}) (1 - p_{34}) \le 0,
\]
so immediate maximality is never better in that example. The paper further states that with no limit on the number of observations, there is a maximum-expectation matching with one edge chosen per observation [1605.08616].

This suggests a broader interpretation of under-matching. In the patience-imbalance model, the planner under-matches to preserve scarce high-quality supply. In the recourse model, the planner may under-match to preserve feasible future rearrangements after uncertainty resolves. The common structure is option preservation: matching fewer pairs now can be optimal because it improves the continuation problem.

## 5. Under-matching and latent compatibility in learned matching

A related, though not identical, issue arises in "Learning to Match" [1707.09678], which studies repeated assignment of workers to tasks when worker skills are latent and task requirements are known. Here the platform repeatedly chooses assignments
\[
a_k \in A,
\]
observes noisy rewards
\[
r_k \sim P(a,\skill,\Tasks),
\]
and seeks to maximize cumulative reward
\[
\sum_k r_k.
\]
Workers have unknown multidimensional skill vectors, tasks have known requirement vectors, and observed feedback is a coarse success/failure signal rather than a direct observation of every edge value [1707.09678].

The paper’s main algorithmic contribution is Hungarian min-max estimation (HME). For each worker-skill coordinate it maintains lower and upper bounds \(s^{\min}_{\worker,i}\) and \(s^{\max}_{\worker,i}\), updates them after observing binary task outcomes, forms point estimates
\[
\widehat{\skill}_{\worker,i} = \dfrac{s^{\min}_{i} + s^{\max}_{i}{2},
\]
and then solves the current assignment problem using the Hungarian algorithm on estimated rewards. In synthetic experiments with \(10\) workers, up to \(300\) tasks, \(3\) skills/requirements, and noise level around \(0.15\), HME reaches roughly \(85\%\) to \(87\%\) of oracle performance [1707.09678].

This literature does not define under-matching in the dynamic-rationing sense. A plausible implication, however, is that when compatibility is latent, aggressive immediate matching can be myopic for a different reason: it may exploit current estimates without generating the information needed for future assignment quality. The paper is explicit that HME handles exploration only implicitly and that explicit optimism or randomized exploration is left for future work. In that sense, matching under uncertainty introduces an epistemic analogue of under-matching: the platform may need to withhold, diversify, or otherwise reshape immediate assignments because learning itself has continuation value.

## 6. Welfare comparisons and institutional implications

The welfare consequences of under-matching depend sharply on market architecture. In the centralized patience-imbalance model, the paper compares three systems under symmetric arrivals \(p=q\): full backlog on both sides, one-sided backlog, and no backlog. Centralized welfare is weakly ordered as
\[
W^{ce}_{FB}\ge W^{ce}_{OB}\ge W^{ce}_{NB},
\]
with one-sided backlog lying between full patience and complete impatience. This means patience strictly enlarges the scope for efficient selective matching, and one-sided patience captures part, but not all, of the gains from full intertemporal coordination [2602.03995].

In decentralized systems, the ranking is not monotone. The paper shows that there exist thresholds \(\alpha_1<\alpha_2\) such that low payoff shares to supply yield
\[
W^{de}_{FB}\ge W^{de}_{OB}\ge W^{de}_{NB},
\]
high payoff shares can reverse the order,
\[
W^{de}_{FB}\le W^{de}_{OB}\le W^{de}_{NB},
\]
and for intermediate \(\alpha\) the ordering depends on \(h\). Thus enabling patience can either increase or decrease welfare once selective waiting is privately chosen rather than centrally rationed [2602.03995].

The institutional implication is not that under-matching is universally desirable or undesirable. Its welfare status depends on whether it is generated by socially aligned rationing, by strategic holdout behavior, or by the need to preserve future flexibility or information. The literature summarized here supports a general conclusion: the naive principle “execute every feasible match immediately” is not robust once continuation values matter. Those continuation values may arise from assortative gains under patience imbalance, from residual-graph optionality under recourse, or from learning about latent match quality. The analytical task is therefore not to eliminate under-matching categorically, but to characterize when it is the efficient response to dynamic constraints and when it is a decentralized distortion requiring design or incentive correction.

Source: https://www.emergentmind.com/topics/under-matching