---
title: Constituent Subtraction
url: https://www.emergentmind.com/topics/constituent-subtraction
type: topic
---

# Constituent Subtraction

Searching arXiv for papers on “constituent subtraction” across the senses represented in the provided data.
Constituent subtraction is a polysemous technical term used in at least three distinct research contexts. In convex analysis, it denotes a generalized Minkowski–Pontryagin difference between compact convex sets, formalized as a collection of inclusion-minimal convex addends and endowed with a linear vector-space structure [1801.06946]. In high-energy physics, it denotes a local background-removal procedure for jets and event observables, implemented by matching physical particles to uniformly distributed “ghost” particles that encode pileup or underlying-event contamination; its iterative extension, Iterative Constituent Subtraction, applies the procedure event-wide and redistributes residual background over multiple passes [1905.03470]. In cognitive science and computational linguistics, constituent subtraction—also called constituent deletion—denotes deletion of a contiguous span of words that exactly corresponds to a node in a latent constituency tree, and is used as an out-of-distribution behavioral probe of hierarchical sentence representations in humans and large language models [2405.18241]. The shared label does not indicate a common formalism across these domains; rather, each usage concerns structured removal relative to an underlying compositional representation.

## 1. Convex-set constituent subtraction as generalized Minkowski–Pontryagin difference

Nurminski and Uryasev define constituent subtraction on the family $\K$ of all nonempty compact convex subsets of a finite-dimensional Euclidean space $E$ by
\[
X\ominus Y
\;=\;
\bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X,\quad Z\text{ is minimal (by set-inclusion) among all }W\text{ with }Y+W\subseteq X\bigr\}.
\]
Equivalently,
\[
X\ominus Y
=
\bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X\text{ and there is no }W\subsetneq Z\text{ with }Y+W\subseteq X\bigr\}.
\]
Here the “difference” is not generally a single set, but a collection of minimal convex sets whose Minkowski sum with $Y$ remains inside $X$ [1801.06946].

The construction is motivated by the desire to invert Minkowski addition without restricting attention to cases where a classical Minkowski–Pontryagin difference exists as a single convex set. The definition is therefore set-valued at a higher level: $X\ominus Y$ is treated as an element of a new space of differences. This suggests that the operation is intended less as a subtraction of points and more as a representation of all irreducible convex addends compatible with the inclusion constraint.

## 2. Linear structure and exact inversion properties

The same work regards each difference $X\ominus Y$ as a “point” in a new space
\[
{\mathcal D}=\{\,X\ominus Y: X,Y\in\K\}.
\]
On this family, scalar multiplication and addition are defined by
\[
\alpha\,(X\ominus Y)
:=(\,\alpha X)\;\ominus\;(\,\alpha Y),
\]
and
\[
(X\ominus Y)\;+\;(Z\ominus W)
:=
(\,X+Z)\;\ominus\;(Y+W).
\]
With these operations, the usual axioms of a real vector space are checked: commutativity and associativity of “$+$” follow from those of Minkowski sum; the zero element is $0\in\K$ because $X\ominus X=0$ and $0\ominus0=0$; and additive inverses satisfy
\[
-(X\ominus Y)=(Y\ominus X)
\]
[1801.06946].

A central property is cancellation:
\[
X\ominus Y=(X+Z)\ominus(Y+Z)
\]
for any $Z\in\K$. In particular, taking $Z=Y$ yields
\[
(X\ominus Y)+Y
=\;(X+Y)\;\ominus\;(Y+Y)
=\;X\;\ominus\;0
=\;X.
\]
The paper states that Minkowski summation by $Y$ is exactly inverted by adding $Y$ to the difference. It also gives the invertibility lemma
\[
X\ominus Y=0 \quad \text{if and only if} \quad X=Y
\]
[1801.06946].

These statements distinguish the construction from partial difference operations that recover subtraction only under restrictive containment hypotheses. A plausible implication is that the vector-space formalization is meant to supply an algebraic environment in which convex-set manipulations can be performed with the same formal convenience as linear operations on vectors.

## 3. Support functions, $\epsilon$-subdifferentials, and worked examples

For any convex compact set $A$, the support function is written as
\[
(A)_p=\sup\{\,p\cdot a:a\in A\}.
\]
The key identity proved for the subtraction is
\[
\inf_{Z\in X\ominus Y}(Z)_p
\;=\;
(X)_p\;-\;(Y)_p
\]
for all $X,Y\in\K$ and all $p\in E$ [1801.06946]. In words, the pointwise infimum of the support functions of all minimal addends recovers the usual difference of support functions.

The main application given is a proof of Lipschitz continuity of $\epsilon$-subdifferentials. For a proper closed convex function $f:E\to\R\cup\{+\infty\}$, fixed $x\in\dom f$, and
\[
D(\epsilon)=\partial_\epsilon f(x)
=
\{\,g\in E:\;f(y)\ge f(x)+g\cdot(y-x)-\epsilon\;\forall y\},
\]
the paper notes that for each $\epsilon\ge0$, $D(\epsilon)\in\K$, that $D(0)=\partial f(x)$, and that $D(\epsilon)$ grows in $\epsilon$. The theorem states that for any fixed $\bar\epsilon>0$ and any $\upsilon\in(0,\bar\epsilon)$, there is a constant $L$ such that for all $\epsilon',\epsilon''\in[\bar\epsilon-\upsilon,\bar\epsilon+\upsilon]$,
\[
d_H\bigl(D(\epsilon'),D(\epsilon'')\bigr)
\;\le\;
L\,\bigl|\epsilon'-\epsilon''\bigr|,
\]
where $d_H$ is the Hausdorff distance [1801.06946]. The proof sketch uses convexity of the graph of $\epsilon\mapsto D(\epsilon)$, rearrangement, and the subtraction $D(\epsilon'')\ominus D(0)$ to derive a linear bound.

Several concrete examples are provided. For intervals in $\R$, if $X=[a,b]$ and $Y=[c,d]$, then the unique minimal $Z$ is $[a-d,b-c]$, so
\[
[a,b]\ominus[c,d]
=
\bigl\{[a-d,b-c]\bigr\}.
\]
For a triangle $A$ with vertices $M=(0,0),N=(1,0),N'=(0,1)$ and its base segment $B=[M,N]$, the difference consists of all minimal line segments $Z$ satisfying $B+Z\subset A$. For Euclidean discs of radii $r_X\ge r_Y$ centered at the origin,
\[
X\ominus Y
=
\{\,\text{the single disc of radius }r_X-r_Y\}
\]
[1801.06946].

## 4. Constituent subtraction in jet physics

In collider phenomenology, Constituent Subtraction refers to a per-particle background-removal algorithm for jets and event observables such as missing transverse energy. The method addresses contamination from soft background coming from pileup in proton-proton collisions, or underlying event in heavy-ion collisions [1905.03470].

The formulation begins with estimation of the background transverse-momentum density. As in the area–median method, one may partition the event into patches or use ghost jets. A common estimator is
\[
\rho \;=\;\mathrm{median}_{i}\Bigl\{\frac{p_{T,i}}{A_i}\Bigr\},
\]
possibly as a function of rapidity, $\rho=\rho(y)$, to account for non-uniformities. With a fixed grid of spacing $\Delta y\times\Delta\phi$, the local estimate is written
\[
\rho(y)\;=\;\mathrm{median}_{j\,:|y_j-y|<\Delta y/2}\Bigl\{\frac{p_{T,j}}{A_{\rm cell}}\Bigr\}
\]
[1905.03470].

The algorithm introduces a dense set of “ghost” particles, each of area $A_{\rm ghost}$, uniformly covering $|y|<y_{\max}$ and initially assigned
\[
p_{T}^{(g)}=\rho\,A_{\rm ghost}.
\]
For each real particle $i$ and ghost $k$, it defines the distance
\[
d_{ik}=p_{T,i}^{\;\alpha}\,\Delta R_{ik},
\qquad
\Delta R_{ik}=\sqrt{(y_i-y_k)^2+(\phi_i-\phi_k)^2},
\]
where $\alpha$ is a free parameter. All pairs are sorted by increasing $d_{ik}$, and the algorithm iterates through the list with the update
\[
\text{if }p_{T,i}\ge p_{T,k}:\quad
p_{T,i}\leftarrow p_{T,i}-p_{T,k},\;\;
p_{T,k}\leftarrow 0;
\quad
\text{else}:\;
p_{T,k}\leftarrow p_{T,k}-p_{T,i},\;\;
p_{T,i}\leftarrow 0.
\]
The procedure stops whenever the current pair has $\Delta R_{ik}>\Delta R_{\max}$, and the surviving real particles with reduced $p_T$ are used for clustering or observable computation [1905.03470].

This formulation is fully local and constituent-level. Unlike jet-area corrections that act primarily at the jet level, the subtraction modifies the constituents themselves, which permits simultaneous correction of jet kinematics and substructure observables.

## 5. Iterative Constituent Subtraction and performance characteristics

Iterative Constituent Subtraction (ICS) extends event-wide Constituent Subtraction by redistributing residual ghost $p_T$ that remains after a finite-$\Delta R_{\max}$ pass and repeating the procedure. The rationale given is that event-wide CS with a finite $\Delta R_{\max}$ tends to leave some residual ghost $p_T$ un-subtracted, especially if $\Delta R_{\max}$ is small; ICS “equilibrates” the background subtraction across the entire event [1905.03470].

Its algorithmic steps are: estimate $\rho$; initialize ghosts $G_{\rm in}$ with $p_T^g=\rho\,A_{\rm ghost}$; for $n=1\ldots N_{\rm iter}$ run event-wide CS on particle set $P$ and ghost set $G_{\rm in}$ with parameters $(\alpha,\Delta R_{\max}^{(n)},A_{\rm ghost})$; compute
\[
p_T^{\rm in}=\sum_{g\in G_{\rm in}} p_T^g,
\qquad
p_T^{\rm out}=\sum_{g\in G_{\rm out}} p_T^g;
\]
optionally remove residual ghosts; and rescale surviving input ghosts by
\[
p_T^g\leftarrow p_T^g\times\frac{p_T^{\rm out}}{p_T^{\rm in}}.
\]
The final corrected particle set $P$ is then returned. Convergence is usually controlled by a fixed $N_{\rm iter}$, and the paper states that in practice 2 or 3 iterations are sufficient because gains beyond that are marginal [1905.03470].

The same work provides empirical parameter recommendations. For event-wide CS, it gives $A_{\rm ghost}=0.0025$, $\alpha=1$, and $\Delta R_{\max}=0.25$ for $R=0.4$ jets or $0.7$ for $R=1.0$ jets. For ICS with two iterations and ghost removal, it gives $\alpha=1$, $A_{\rm ghost}=0.0025$, $(\Delta R_{\max}^{(1)},\Delta R_{\max}^{(2)})=(0.20,0.10)$ for $R=0.4$, and $(0.20,0.35)$ for $R=1.0$ [1905.03470].

Performance is evaluated using Jet Energy Scale,
\[
\langle p_{T,\rm reco}/p_{T,\rm true}\rangle,
\]
Jet Energy Resolution,
\[
\sigma(p_{T,\rm reco}/p_{T,\rm true}),
\]
with
\[
\mathrm{JER}(p_T)=\frac{a}{\sqrt{p_T}}\oplus\frac{b}{p_T}\oplus c,
\]
and shape-observable bias and resolution,
\[
\mathrm{bias}(x)=\frac{\langle x^{\rm rec}-x^{\rm true}\rangle}{\langle x^{\rm true}\rangle},
\qquad
\mathrm{res}(x)=\frac{\mathrm{RMS}(x^{\rm rec}-x^{\rm true})}{\langle x^{\rm true}\rangle}.
\]
Compared to jet-by-jet CS, ICS reduces the JER noise term by up to 30% at high pileup, and a further 5–10% relative to event-wide CS. The bias on $p_T,m,\mathrm{width},\tau_{21},\tau_{32}$ is as small or smaller than competing methods, and ICS typically yields the best resolution–bias trade-off across $R=0.4$ and $R=1.0$ jets over pileup $0<\mu<140$. For missing $E_T$, ICS or event-wide CS+SoftKiller reduce the RMS of $E_{T,x}^{\rm miss}$ by $\sim20$% at large $\mu$ compared to SoftKiller alone [1905.03470].

A representative comparison for $R=1.0$, 250–300 GeV jets at $\mu=100$ is summarized below.

| Algorithm | JES bias | JER ($\sigma_{p_T}$) |
|---|---:|---:|
| Area–median | $\simeq 0$% | 1.00 |
| Jet-by-jet CS | $\simeq 0$% | 0.90 |
| Event-wide CS | $\simeq 0$% | 0.75 |
| ICS (2 iter) | $\simeq 0$% | 0.68 |

Implementation details noted in the paper include FastJet Contrib support, computational scaling roughly like $O(N_{\rm ghosts}\times N_{\rm particles})$ plus sorting, a practical ghost-area range $A_{\rm ghost}\approx0.0025$–$0.01$, the need for consistent $\max\eta$ choices in $\rho(y)$ estimation and ghost placement, and parameter retuning in very high-density environments [1905.03470].

## 6. Constituent subtraction as linguistic constituent deletion

In cognitive science and NLP-oriented work, constituent subtraction means deletion of a contiguous span of words from a sentence that exactly corresponds to a node in the sentence’s latent constituency tree, such as a complete noun phrase or verb phrase. It is contrasted with non-constituent deletion, which removes a contiguous word string that does not form a well-formed constituent in the tree [2405.18241].

The paper studies a one-shot word-deletion task. Participants see a single demonstration in which a constituent has been deleted and then apply the same transformation to a novel test sentence. Deleting a constituent in the test is treated as a signature of applying an underlying constituency-based rule, whereas deleting a non-constituent suggests reliance on surface cues such as word order rather than hierarchical structure [2405.18241].

The experimental design includes English sentences from the Penn Treebank and Chinese sentences from the Chinese Treebank. Demonstrations in Experiment 1a targeted NP directly under VP, while Experiment 2 allowed any constituent. Test sentences contained at least one constituent of the same node category and at least one constituent with the same parent category, so that node-category and parent-category rules could be dissociated. Additional experiments used parallel Chinese–English sentences with identical tree shapes, full-tree reconstruction targets, and syntactically ambiguous sentences with one semantically plausible versus implausible parse [2405.18241].

Humans comprised 30 native speakers per language per experiment, plus 30 high-proficiency L2 English learners in selected conditions. The LLM condition included ChatGPT with temperature $=0$ and max\_tokens $=200$, and in Experiment 2 also GPT-4, Claude-3, Gemini-1, and Llama-3. A naïve LSTM baseline with static word embeddings and sinusoidal positional embeddings was trained only on the deletion task to test whether linear-sequence cues suffice [2405.18241].

Behavioral measures include constituent rate
\[
CR = \frac{\#\text{ deletions that form complete constituent}}{\#\text{ valid tests}},
\]
non-constituent rate, “other” rate, explained ratio for a rule $R$,
\[
ER_R = \frac{\#\text{ deletions consistent with }R}{\#\text{ constituent deletions}},
\]
and, in tree reconstruction,
\[
ER(T)=\frac{\sum_{deleted\ spans\ s} \mathbf{1}[s\in Nodes(T)]}{\text{total \# deletions}}.
\]
The paper also defines tree overlap with the gold tree by
\[
F1 = \frac{2\cdot|Nodes(\hat T)\cap Nodes(T_{gold})|}{|Nodes(\hat T)|+|Nodes(T_{gold})|}
\]
and balance factor as the number of right-branch edges divided by the number of left-branch edges [2405.18241].

## 7. Behavioral findings, CKY reconstruction, and interpretive scope

The quantitative findings are organized around four experiments. In Experiment 1, humans achieved constituent rate $\approx 0.93$ in English and $0.97$ in Chinese, both with $p<0.001$ versus chance; ChatGPT achieved $\approx 0.87$ in English and $0.43$ in Chinese, both with $p<0.001$; “other” responses were below 5%; and the LSTM one-shot baseline was at chance, with $CR\approx0.52$ in English and $0.50$ in Chinese, not above chance and needing about 50 demonstrations to match human performance. Rule inference showed that in English both humans and ChatGPT strongly preferred the parent-category rule, with $ER_{parent}\approx0.75$ versus $ER_{node}\approx0.25$, whereas in Chinese both preferred the node-category rule; L2 English learners switched to the parent rule when prompted in English; and parallel Chinese–English sentences confirmed that rule preference depends on input language rather than sentence shape [2405.18241].

Experiment 2 reported similar constituent rates, with humans at approximately $0.90$ and ChatGPT at approximately $0.87$, both above chance, while the LSTM remained at chance in the one-shot setting and required roughly 10–50 shots to match humans or ChatGPT. Experiment 3 used the deletions to reconstruct trees: the deletion-based tree explained about 80% of human and ChatGPT deletions, with $p<0.001$ versus random trees; F1 overlap with the gold tree was about $0.60$, above chance of about $0.40$; average depth and width were similar to linguistic trees; and English trees, both gold and reconstructed, were significantly more right-branching than Chinese. Experiment 4 found that native English speakers showed no significant difference between deleting semantically plausible and implausible spans, L2 learners showed a small semantic effect, and ChatGPT showed a weak but significant reverse semantic bias; on syntactic nonsense sentences, ChatGPT still deleted constituents above chance [2405.18241].

For tree reconstruction, the paper formulates a CKY-style dynamic program. Given a sentence $S$ of length $L$ and deleted spans $D=\{s_j\}$ with $s_j=(i_j,k_j)$, it seeks a binary constituency tree $T$ maximizing
\[
ER(T)=\frac{1}{|D|}\sum_{s\in D}\mathbf{1}[s\text{ is a node span in }T].
\]
Using span scores from deletion frequencies, the recurrence is
\[
V[i,i]=\mathrm{score}(i,i),
\]
and for span $(i,k)$ with $k>i$,
\[
V[i,k]=\max\Bigl\{\mathrm{score}(i,k),\ \max_{i<m<k}(V[i,m]+V[m,k])\Bigr\},
\]
with the tree recovered by backtracking the maximizing splits [2405.18241].

The broader interpretation offered in the paper is that both humans and LLMs actively use latent hierarchical constituency representations even in an unfamiliar one-shot deletion task that contains no explicit linguistic instruction, whereas a naïve LSTM with only word- and position-level cues fails one-shot. It further states that constituent subtraction can serve as a general, out-of-distribution probe for latent syntax in closed-source models and that the results bridge Marr’s levels of analysis by suggesting convergence at the algorithmic level despite differing implementations [2405.18241]. This suggests that, in this domain, “constituent subtraction” functions less as an editing operation than as an experimental assay of internal structure.

Source: https://www.emergentmind.com/topics/constituent-subtraction