Constituent Subtraction
- Constituent subtraction is a polysemous concept describing the structured removal of parts in convex sets, jets, and sentences using domain-specific formalism.
- In convex analysis, it is formalized as a generalized Minkowski–Pontryagin difference that recovers subtraction via minimal convex addends and vector-space operations.
- In jet physics and linguistics, iterative constituent subtraction refines background corrections and probes latent hierarchical structures, enhancing both performance and interpretation.
Searching arXiv for papers on “constituent subtraction” across the senses represented in the provided data. Constituent subtraction is a polysemous technical term used in at least three distinct research contexts. In convex analysis, it denotes a generalized Minkowski–Pontryagin difference between compact convex sets, formalized as a collection of inclusion-minimal convex addends and endowed with a linear vector-space structure (Nurminski et al., 2018). In high-energy physics, it denotes a local background-removal procedure for jets and event observables, implemented by matching physical particles to uniformly distributed “ghost” particles that encode pileup or underlying-event contamination; its iterative extension, Iterative Constituent Subtraction, applies the procedure event-wide and redistributes residual background over multiple passes (Berta et al., 2019). In cognitive science and computational linguistics, constituent subtraction—also called constituent deletion—denotes deletion of a contiguous span of words that exactly corresponds to a node in a latent constituency tree, and is used as an out-of-distribution behavioral probe of hierarchical sentence representations in humans and LLMs (Liu et al., 2024). The shared label does not indicate a common formalism across these domains; rather, each usage concerns structured removal relative to an underlying compositional representation.
1. Convex-set constituent subtraction as generalized Minkowski–Pontryagin difference
Nurminski and Uryasev define constituent subtraction on the family $\K$ of all nonempty compact convex subsets of a finite-dimensional Euclidean space by
$X\ominus Y \;=\; \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X,\quad Z\text{ is minimal (by set-inclusion) among all }W\text{ with }Y+W\subseteq X\bigr\}.$
Equivalently,
$X\ominus Y = \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X\text{ and there is no }W\subsetneq Z\text{ with }Y+W\subseteq X\bigr\}.$
Here the “difference” is not generally a single set, but a collection of minimal convex sets whose Minkowski sum with remains inside (Nurminski et al., 2018).
The construction is motivated by the desire to invert Minkowski addition without restricting attention to cases where a classical Minkowski–Pontryagin difference exists as a single convex set. The definition is therefore set-valued at a higher level: is treated as an element of a new space of differences. This suggests that the operation is intended less as a subtraction of points and more as a representation of all irreducible convex addends compatible with the inclusion constraint.
2. Linear structure and exact inversion properties
The same work regards each difference as a “point” in a new space
${\mathcal D}=\{\,X\ominus Y: X,Y\in\K\}.$
On this family, scalar multiplication and addition are defined by
and
0
With these operations, the usual axioms of a real vector space are checked: commutativity and associativity of “1” follow from those of Minkowski sum; the zero element is 2 because 3 and 4; and additive inverses satisfy
5
A central property is cancellation: 6 for any 7. In particular, taking 8 yields
9
The paper states that Minkowski summation by $X\ominus Y \;=\; \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X,\quad Z\text{ is minimal (by set-inclusion) among all }W\text{ with }Y+W\subseteq X\bigr\}.$0 is exactly inverted by adding $X\ominus Y \;=\; \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X,\quad Z\text{ is minimal (by set-inclusion) among all }W\text{ with }Y+W\subseteq X\bigr\}.$1 to the difference. It also gives the invertibility lemma
$X\ominus Y \;=\; \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X,\quad Z\text{ is minimal (by set-inclusion) among all }W\text{ with }Y+W\subseteq X\bigr\}.$2
These statements distinguish the construction from partial difference operations that recover subtraction only under restrictive containment hypotheses. A plausible implication is that the vector-space formalization is meant to supply an algebraic environment in which convex-set manipulations can be performed with the same formal convenience as linear operations on vectors.
3. Support functions, $X\ominus Y \;=\; \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X,\quad Z\text{ is minimal (by set-inclusion) among all }W\text{ with }Y+W\subseteq X\bigr\}.$3-subdifferentials, and worked examples
For any convex compact set $X\ominus Y \;=\; \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X,\quad Z\text{ is minimal (by set-inclusion) among all }W\text{ with }Y+W\subseteq X\bigr\}.$4, the support function is written as
$X\ominus Y \;=\; \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X,\quad Z\text{ is minimal (by set-inclusion) among all }W\text{ with }Y+W\subseteq X\bigr\}.$5
The key identity proved for the subtraction is
$X\ominus Y \;=\; \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X,\quad Z\text{ is minimal (by set-inclusion) among all }W\text{ with }Y+W\subseteq X\bigr\}.$6
for all $X\ominus Y \;=\; \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X,\quad Z\text{ is minimal (by set-inclusion) among all }W\text{ with }Y+W\subseteq X\bigr\}.$7 and all $X\ominus Y \;=\; \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X,\quad Z\text{ is minimal (by set-inclusion) among all }W\text{ with }Y+W\subseteq X\bigr\}.$8 (Nurminski et al., 2018). In words, the pointwise infimum of the support functions of all minimal addends recovers the usual difference of support functions.
The main application given is a proof of Lipschitz continuity of $X\ominus Y \;=\; \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X,\quad Z\text{ is minimal (by set-inclusion) among all }W\text{ with }Y+W\subseteq X\bigr\}.$9-subdifferentials. For a proper closed convex function $X\ominus Y = \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X\text{ and there is no }W\subsetneq Z\text{ with }Y+W\subseteq X\bigr\}.$0, fixed $X\ominus Y = \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X\text{ and there is no }W\subsetneq Z\text{ with }Y+W\subseteq X\bigr\}.$1, and
$X\ominus Y = \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X\text{ and there is no }W\subsetneq Z\text{ with }Y+W\subseteq X\bigr\}.$2
the paper notes that for each $X\ominus Y = \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X\text{ and there is no }W\subsetneq Z\text{ with }Y+W\subseteq X\bigr\}.$3, $X\ominus Y = \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X\text{ and there is no }W\subsetneq Z\text{ with }Y+W\subseteq X\bigr\}.$4, that $X\ominus Y = \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X\text{ and there is no }W\subsetneq Z\text{ with }Y+W\subseteq X\bigr\}.$5, and that $X\ominus Y = \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X\text{ and there is no }W\subsetneq Z\text{ with }Y+W\subseteq X\bigr\}.$6 grows in $X\ominus Y = \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X\text{ and there is no }W\subsetneq Z\text{ with }Y+W\subseteq X\bigr\}.$7. The theorem states that for any fixed $X\ominus Y = \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X\text{ and there is no }W\subsetneq Z\text{ with }Y+W\subseteq X\bigr\}.$8 and any $X\ominus Y = \bigl\{\,Z\in\K\;\big|\;Y+Z\subseteq X\text{ and there is no }W\subsetneq Z\text{ with }Y+W\subseteq X\bigr\}.$9, there is a constant 0 such that for all 1,
2
where 3 is the Hausdorff distance (Nurminski et al., 2018). The proof sketch uses convexity of the graph of 4, rearrangement, and the subtraction 5 to derive a linear bound.
Several concrete examples are provided. For intervals in 6, if 7 and 8, then the unique minimal 9 is 0, so
1
For a triangle 2 with vertices 3 and its base segment 4, the difference consists of all minimal line segments 5 satisfying 6. For Euclidean discs of radii 7 centered at the origin,
8
4. Constituent subtraction in jet physics
In collider phenomenology, Constituent Subtraction refers to a per-particle background-removal algorithm for jets and event observables such as missing transverse energy. The method addresses contamination from soft background coming from pileup in proton-proton collisions, or underlying event in heavy-ion collisions (Berta et al., 2019).
The formulation begins with estimation of the background transverse-momentum density. As in the area–median method, one may partition the event into patches or use ghost jets. A common estimator is
9
possibly as a function of rapidity, 0, to account for non-uniformities. With a fixed grid of spacing 1, the local estimate is written
2
The algorithm introduces a dense set of “ghost” particles, each of area 3, uniformly covering 4 and initially assigned
5
For each real particle 6 and ghost 7, it defines the distance
8
where 9 is a free parameter. All pairs are sorted by increasing 0, and the algorithm iterates through the list with the update
1
The procedure stops whenever the current pair has 2, and the surviving real particles with reduced 3 are used for clustering or observable computation (Berta et al., 2019).
This formulation is fully local and constituent-level. Unlike jet-area corrections that act primarily at the jet level, the subtraction modifies the constituents themselves, which permits simultaneous correction of jet kinematics and substructure observables.
5. Iterative Constituent Subtraction and performance characteristics
Iterative Constituent Subtraction (ICS) extends event-wide Constituent Subtraction by redistributing residual ghost 4 that remains after a finite-5 pass and repeating the procedure. The rationale given is that event-wide CS with a finite 6 tends to leave some residual ghost 7 un-subtracted, especially if 8 is small; ICS “equilibrates” the background subtraction across the entire event (Berta et al., 2019).
Its algorithmic steps are: estimate 9; initialize ghosts ${\mathcal D}=\{\,X\ominus Y: X,Y\in\K\}.$0 with ${\mathcal D}=\{\,X\ominus Y: X,Y\in\K\}.$1; for ${\mathcal D}=\{\,X\ominus Y: X,Y\in\K\}.$2 run event-wide CS on particle set ${\mathcal D}=\{\,X\ominus Y: X,Y\in\K\}.$3 and ghost set ${\mathcal D}=\{\,X\ominus Y: X,Y\in\K\}.$4 with parameters ${\mathcal D}=\{\,X\ominus Y: X,Y\in\K\}.$5; compute
${\mathcal D}=\{\,X\ominus Y: X,Y\in\K\}.$6
optionally remove residual ghosts; and rescale surviving input ghosts by
${\mathcal D}=\{\,X\ominus Y: X,Y\in\K\}.$7
The final corrected particle set ${\mathcal D}=\{\,X\ominus Y: X,Y\in\K\}.$8 is then returned. Convergence is usually controlled by a fixed ${\mathcal D}=\{\,X\ominus Y: X,Y\in\K\}.$9, and the paper states that in practice 2 or 3 iterations are sufficient because gains beyond that are marginal (Berta et al., 2019).
The same work provides empirical parameter recommendations. For event-wide CS, it gives 0, 1, and 2 for 3 jets or 4 for 5 jets. For ICS with two iterations and ghost removal, it gives 6, 7, 8 for 9, and 00 for 01 (Berta et al., 2019).
Performance is evaluated using Jet Energy Scale,
02
Jet Energy Resolution,
03
with
04
and shape-observable bias and resolution,
05
Compared to jet-by-jet CS, ICS reduces the JER noise term by up to 30% at high pileup, and a further 5–10% relative to event-wide CS. The bias on 06 is as small or smaller than competing methods, and ICS typically yields the best resolution–bias trade-off across 07 and 08 jets over pileup 09. For missing 10, ICS or event-wide CS+SoftKiller reduce the RMS of 11 by 12% at large 13 compared to SoftKiller alone (Berta et al., 2019).
A representative comparison for 14, 250–300 GeV jets at 15 is summarized below.
| Algorithm | JES bias | JER (16) |
|---|---|---|
| Area–median | 17% | 1.00 |
| Jet-by-jet CS | 18% | 0.90 |
| Event-wide CS | 19% | 0.75 |
| ICS (2 iter) | 20% | 0.68 |
Implementation details noted in the paper include FastJet Contrib support, computational scaling roughly like 21 plus sorting, a practical ghost-area range 22–23, the need for consistent 24 choices in 25 estimation and ghost placement, and parameter retuning in very high-density environments (Berta et al., 2019).
6. Constituent subtraction as linguistic constituent deletion
In cognitive science and NLP-oriented work, constituent subtraction means deletion of a contiguous span of words from a sentence that exactly corresponds to a node in the sentence’s latent constituency tree, such as a complete noun phrase or verb phrase. It is contrasted with non-constituent deletion, which removes a contiguous word string that does not form a well-formed constituent in the tree (Liu et al., 2024).
The paper studies a one-shot word-deletion task. Participants see a single demonstration in which a constituent has been deleted and then apply the same transformation to a novel test sentence. Deleting a constituent in the test is treated as a signature of applying an underlying constituency-based rule, whereas deleting a non-constituent suggests reliance on surface cues such as word order rather than hierarchical structure (Liu et al., 2024).
The experimental design includes English sentences from the Penn Treebank and Chinese sentences from the Chinese Treebank. Demonstrations in Experiment 1a targeted NP directly under VP, while Experiment 2 allowed any constituent. Test sentences contained at least one constituent of the same node category and at least one constituent with the same parent category, so that node-category and parent-category rules could be dissociated. Additional experiments used parallel Chinese–English sentences with identical tree shapes, full-tree reconstruction targets, and syntactically ambiguous sentences with one semantically plausible versus implausible parse (Liu et al., 2024).
Humans comprised 30 native speakers per language per experiment, plus 30 high-proficiency L2 English learners in selected conditions. The LLM condition included ChatGPT with temperature 26 and max_tokens 27, and in Experiment 2 also GPT-4, Claude-3, Gemini-1, and Llama-3. A naĂŻve LSTM baseline with static word embeddings and sinusoidal positional embeddings was trained only on the deletion task to test whether linear-sequence cues suffice (Liu et al., 2024).
Behavioral measures include constituent rate
28
non-constituent rate, “other” rate, explained ratio for a rule 29,
30
and, in tree reconstruction,
31
The paper also defines tree overlap with the gold tree by
32
and balance factor as the number of right-branch edges divided by the number of left-branch edges (Liu et al., 2024).
7. Behavioral findings, CKY reconstruction, and interpretive scope
The quantitative findings are organized around four experiments. In Experiment 1, humans achieved constituent rate 33 in English and 34 in Chinese, both with 35 versus chance; ChatGPT achieved 36 in English and 37 in Chinese, both with 38; “other” responses were below 5%; and the LSTM one-shot baseline was at chance, with 39 in English and 40 in Chinese, not above chance and needing about 50 demonstrations to match human performance. Rule inference showed that in English both humans and ChatGPT strongly preferred the parent-category rule, with 41 versus 42, whereas in Chinese both preferred the node-category rule; L2 English learners switched to the parent rule when prompted in English; and parallel Chinese–English sentences confirmed that rule preference depends on input language rather than sentence shape (Liu et al., 2024).
Experiment 2 reported similar constituent rates, with humans at approximately 43 and ChatGPT at approximately 44, both above chance, while the LSTM remained at chance in the one-shot setting and required roughly 10–50 shots to match humans or ChatGPT. Experiment 3 used the deletions to reconstruct trees: the deletion-based tree explained about 80% of human and ChatGPT deletions, with 45 versus random trees; F1 overlap with the gold tree was about 46, above chance of about 47; average depth and width were similar to linguistic trees; and English trees, both gold and reconstructed, were significantly more right-branching than Chinese. Experiment 4 found that native English speakers showed no significant difference between deleting semantically plausible and implausible spans, L2 learners showed a small semantic effect, and ChatGPT showed a weak but significant reverse semantic bias; on syntactic nonsense sentences, ChatGPT still deleted constituents above chance (Liu et al., 2024).
For tree reconstruction, the paper formulates a CKY-style dynamic program. Given a sentence 48 of length 49 and deleted spans 50 with 51, it seeks a binary constituency tree 52 maximizing
53
Using span scores from deletion frequencies, the recurrence is
54
and for span 55 with 56,
57
with the tree recovered by backtracking the maximizing splits (Liu et al., 2024).
The broader interpretation offered in the paper is that both humans and LLMs actively use latent hierarchical constituency representations even in an unfamiliar one-shot deletion task that contains no explicit linguistic instruction, whereas a naïve LSTM with only word- and position-level cues fails one-shot. It further states that constituent subtraction can serve as a general, out-of-distribution probe for latent syntax in closed-source models and that the results bridge Marr’s levels of analysis by suggesting convergence at the algorithmic level despite differing implementations (Liu et al., 2024). This suggests that, in this domain, “constituent subtraction” functions less as an editing operation than as an experimental assay of internal structure.