Papers
Topics
Authors
Recent
Search
2000 character limit reached

Non-Overlap Bounds in Theory & Applications

Updated 12 July 2026
  • Non-Overlap Bounds are quantitative statements that limit discrepancies when overlap between entities is absent, with key applications in causal inference, combinatorics, and coding theory.
  • They employ methodologies like uniform likelihood-ratio bounds, linear programming, and Lipschitz extrapolation to derive sharp, domain-specific inequalities and partial-identification intervals.
  • These bounds offer practical tools for error estimation, sensitivity analysis, and feasibility constraints across disciplines, ensuring rigorous assessment of non-overlap scenarios.

Non-overlap bounds are quantitative statements that limit what can occur when overlap, positivity, or compatibility between objects is absent or explicitly restricted. Across the literatures represented here, they appear as partial-identification intervals for causal effects, information-theoretic bounds on discrepancy between treated and control populations, sharp extremal inequalities for words and codes, geometric density limits under controlled intersection, feasibility bounds for scheduling constraints, and Monte Carlo error bounds for near-orthogonal quantum states (D'Amour et al., 2017, Susmann et al., 24 Sep 2025, Zakharov, 23 Feb 2026).

1. Formal notions of overlap and non-overlap

In observational causal inference, overlap is expressed through the propensity score. With binary treatment, population overlap requires 0<e(X1:p)<10 < e(X_{1:p}) < 1 with probability $1$, while strict overlap strengthens this to

η≤e(X1:p)≤1−ηwith probability 1\eta \le e(X_{1:p}) \le 1-\eta \quad \text{with probability 1}

for some fixed η∈(0,0.5)\eta \in (0,0.5). In off-policy evaluation, the analogous support condition is πe(x,a)>0⇒πb(x,a)>0\pi_e(x,a) > 0 \Rightarrow \pi_b(x,a) > 0, and the no-overlap region is Xno={x∈X:πb(x)=0}\mathcal{X}_{\text{no}} = \{x\in\mathcal{X} : \pi_b(x)=0\}. In scheduling, the no-overlap constraint on a single machine is the disjunction ei≤sj∨ej≤sie_i \le s_j \lor e_j \le s_i for every pair of jobs i≠ji\neq j (D'Amour et al., 2017, Khan et al., 2023, Guichard et al., 21 Jan 2026).

In combinatorics on words, overlap is defined through prefix–suffix coincidence. For two words w,u∈Σnw,u\in\Sigma^n, (w,u)(w,u) overlaps if there exists $1$0 such that the suffix of $1$1 of length $1$2 equals the prefix of $1$3 of length $1$4. In coding theory, $1$5 and $1$6 have a $1$7-overlap if the prefix of $1$8 of length $1$9 is identical to the suffix of η≤e(X1:p)≤1−ηwith probability 1\eta \le e(X_{1:p}) \le 1-\eta \quad \text{with probability 1}0 of length η≤e(X1:p)≤1−ηwith probability 1\eta \le e(X_{1:p}) \le 1-\eta \quad \text{with probability 1}1, and a η≤e(X1:p)≤1−ηwith probability 1\eta \le e(X_{1:p}) \le 1-\eta \quad \text{with probability 1}2-overlap-free code forbids such overlaps for every η≤e(X1:p)≤1−ηwith probability 1\eta \le e(X_{1:p}) \le 1-\eta \quad \text{with probability 1}3 with η≤e(X1:p)≤1−ηwith probability 1\eta \le e(X_{1:p}) \le 1-\eta \quad \text{with probability 1}4. These definitions specialize to classical non-overlapping codes when all proper overlaps are forbidden (Zakharov, 23 Feb 2026, Blackburn et al., 2022, Stanovnik, 2024).

In geometric and physical settings, overlap is measured differently but plays the same structural role. In the sphere-packing generalization with limited overlap, one controls either a distance-based overlap

η≤e(X1:p)≤1−ηwith probability 1\eta \le e(X_{1:p}) \le 1-\eta \quad \text{with probability 1}5

or a volume-based overlap

η≤e(X1:p)≤1−ηwith probability 1\eta \le e(X_{1:p}) \le 1-\eta \quad \text{with probability 1}6

For neural quantum states, the overlap is the Hilbert-space inner product η≤e(X1:p)≤1−ηwith probability 1\eta \le e(X_{1:p}) \le 1-\eta \quad \text{with probability 1}7, and non-overlap corresponds to orthogonality or near-orthogonality (Iglesias-Ham et al., 2014, Szołdra, 2023).

2. Overlap as a global restriction in high-dimensional causal inference

A central modern use of non-overlap bounds is to reinterpret strict overlap as a bound on how different treated and control covariate distributions can be. If η≤e(X1:p)≤1−ηwith probability 1\eta \le e(X_{1:p}) \le 1-\eta \quad \text{with probability 1}8 and η≤e(X1:p)≤1−ηwith probability 1\eta \le e(X_{1:p}) \le 1-\eta \quad \text{with probability 1}9 are the conditional covariate distributions under treatment and control, then strict overlap is equivalent to a uniform likelihood-ratio bound

η∈(0,0.5)\eta \in (0,0.5)0

with

η∈(0,0.5)\eta \in (0,0.5)1

This implies bounded η∈(0,0.5)\eta \in (0,0.5)2-divergences, including bounded η∈(0,0.5)\eta \in (0,0.5)3-divergence and bounded Kullback–Leibler divergence between η∈(0,0.5)\eta \in (0,0.5)4 and η∈(0,0.5)\eta \in (0,0.5)5. When η∈(0,0.5)\eta \in (0,0.5)6,

η∈(0,0.5)\eta \in (0,0.5)7

and

η∈(0,0.5)\eta \in (0,0.5)8

These are global non-overlap bounds: if strict overlap holds, treated and control covariate distributions cannot be too far apart in KL or η∈(0,0.5)\eta \in (0,0.5)9 divergence (D'Amour et al., 2017).

The same framework yields explicit mean-imbalance bounds. If πe(x,a)>0⇒πb(x,a)>0\pi_e(x,a) > 0 \Rightarrow \pi_b(x,a) > 00 and πe(x,a)>0⇒πb(x,a)>0\pi_e(x,a) > 0 \Rightarrow \pi_b(x,a) > 01 denote covariate means and πe(x,a)>0⇒πb(x,a)>0\pi_e(x,a) > 0 \Rightarrow \pi_b(x,a) > 02 the covariance matrices, then strict overlap implies

πe(x,a)>0⇒πb(x,a)>0\pi_e(x,a) > 0 \Rightarrow \pi_b(x,a) > 03

Consequently,

πe(x,a)>0⇒πb(x,a)>0\pi_e(x,a) > 0 \Rightarrow \pi_b(x,a) > 04

If the smaller operator norm is πe(x,a)>0⇒πb(x,a)>0\pi_e(x,a) > 0 \Rightarrow \pi_b(x,a) > 05, the average absolute mean difference goes to zero as πe(x,a)>0⇒πb(x,a)>0\pi_e(x,a) > 0 \Rightarrow \pi_b(x,a) > 06. The paper also shows that strict overlap bounds treatment-prediction accuracy: πe(x,a)>0⇒πb(x,a)>0\pi_e(x,a) > 0 \Rightarrow \pi_b(x,a) > 07 for every classifier πe(x,a)>0⇒πb(x,a)>0\pi_e(x,a) > 0 \Rightarrow \pi_b(x,a) > 08. In this sense, strict overlap excludes asymptotically perfect classification and forces the average per-covariate discriminatory information to vanish in high dimension.

3. Partial identification, sensitivity analysis, and structural non-overlap

When overlap fails, one line of work retains the average treatment effect and replaces point identification by partial identification. With bounded outcomes πe(x,a)>0⇒πb(x,a)>0\pi_e(x,a) > 0 \Rightarrow \pi_b(x,a) > 09, the average treatment effect can be decomposed into an identifiable trimmed effect

Xno={x∈X:πb(x)=0}\mathcal{X}_{\text{no}} = \{x\in\mathcal{X} : \pi_b(x)=0\}0

and a worst-case contribution from the non-overlap sets Xno={x∈X:πb(x)=0}\mathcal{X}_{\text{no}} = \{x\in\mathcal{X} : \pi_b(x)=0\}1 and Xno={x∈X:πb(x)=0}\mathcal{X}_{\text{no}} = \{x\in\mathcal{X} : \pi_b(x)=0\}2. This yields the non-overlap bounds

Xno={x∈X:πb(x)=0}\mathcal{X}_{\text{no}} = \{x\in\mathcal{X} : \pi_b(x)=0\}3

so that Xno={x∈X:πb(x)=0}\mathcal{X}_{\text{no}} = \{x\in\mathcal{X} : \pi_b(x)=0\}4. The width is

Xno={x∈X:πb(x)=0}\mathcal{X}_{\text{no}} = \{x\in\mathcal{X} : \pi_b(x)=0\}5

namely the probability mass of the non-overlap population. Because these bounds are non-smooth, the paper introduces smooth approximations Xno={x∈X:πb(x)=0}\mathcal{X}_{\text{no}} = \{x\in\mathcal{X} : \pi_b(x)=0\}6 and Xno={x∈X:πb(x)=0}\mathcal{X}_{\text{no}} = \{x\in\mathcal{X} : \pi_b(x)=0\}7, derives efficient influence functions, proposes a Targeted Minimum Loss-Based estimator, and constructs a multiplier-bootstrap confidence set that is uniformly valid across all trimming and smoothing parameters (Susmann et al., 24 Sep 2025).

A related finite-sample perspective decomposes the full-sample average treatment effect as Xno={x∈X:πb(x)=0}\mathcal{X}_{\text{no}} = \{x\in\mathcal{X} : \pi_b(x)=0\}8, where Xno={x∈X:πb(x)=0}\mathcal{X}_{\text{no}} = \{x\in\mathcal{X} : \pi_b(x)=0\}9 is the contribution from the overlap region ei≤sj∨ej≤sie_i \le s_j \lor e_j \le s_i0 and ei≤sj∨ej≤sie_i \le s_j \lor e_j \le s_i1 is the contribution from the limited-overlap region ei≤sj∨ej≤sie_i \le s_j \lor e_j \le s_i2, with ei≤sj∨ej≤sie_i \le s_j \lor e_j \le s_i3. Under the Lipschitz class

ei≤sj∨ej≤sie_i \le s_j \lor e_j \le s_i4

the paper builds minimax confidence intervals ei≤sj∨ej≤sie_i \le s_j \lor e_j \le s_i5 for ei≤sj∨ej≤sie_i \le s_j \lor e_j \le s_i6. The worst-case bias and variance are controlled through the modulus of continuity ei≤sj∨ej≤sie_i \le s_j \lor e_j \le s_i7, and a combined interval

ei≤sj∨ej≤sie_i \le s_j \lor e_j \le s_i8

is formed by adding an asymptotic interval for ei≤sj∨ej≤sie_i \le s_j \lor e_j \le s_i9 to the minimax interval for i≠ji\neq j0. This converts limited overlap into an explicit sensitivity analysis over the smoothness constant i≠ji\neq j1 (Ma et al., 27 Nov 2025).

Another strand allows simultaneous violations of unconfoundedness and overlap. Under the robust marginal sensitivity model,

i≠ji\neq j2

only a fraction i≠ji\neq j3 of units in group i≠ji\neq j4 satisfy the odds-ratio bound; the remainder may have arbitrary confounding and i≠ji\neq j5 or i≠ji\neq j6. The target estimand is the overlap-weighted average treatment effect

i≠ji\neq j7

which remains identifiable because units with i≠ji\neq j8 or i≠ji\neq j9 receive zero weight. Worst-case lower and upper bounds are computed by linear programming or mixed-integer linear programming under robust sensitivity and covariate-balance constraints, and inference is based on an augmented percentile bootstrap (Cui et al., 16 Sep 2025).

Two further contributions situate non-overlap within explicit modeling architectures. In the Bayesian BART+SPL framework, the sample is split into a region of overlap w,u∈Σnw,u\in\Sigma^n0 and a region of non-overlap w,u∈Σnw,u\in\Sigma^n1 according to the estimated propensity score; Bayesian Additive Regression Trees are used in w,u∈Σnw,u\in\Sigma^n2, while a spline model extrapolates into w,u∈Σnw,u\in\Sigma^n3 with variance inflation

w,u∈Σnw,u\in\Sigma^n4

where w,u∈Σnw,u\in\Sigma^n5 is the distance into non-overlap and w,u∈Σnw,u\in\Sigma^n6 is the range of causal effects in w,u∈Σnw,u\in\Sigma^n7 (Nethery et al., 2018). In fusion problems with randomized and observational data, structural non-overlap induces a feasibility gap

w,u∈Σnw,u\in\Sigma^n8

and under conditional structural non-overlap one obtains

w,u∈Σnw,u\in\Sigma^n9

Representation learning can reduce this gap by recovering overlap in a latent space, but the resulting oracle inequality still decomposes excess risk into overlap recovery, moment violation, and statistical terms (Du et al., 26 Feb 2026).

4. Off-policy evaluation beyond overlap

In off-policy evaluation, the value of a target policy is

(w,u)(w,u)0

and standard importance-weighting methods require the support condition (w,u)(w,u)1. Without this condition, the paper decomposes the value into an identified overlap part and an unidentified no-overlap part,

(w,u)(w,u)2

then partially identifies (w,u)(w,u)3 under the Lipschitz class

(w,u)(w,u)4

possibly intersected with boundedness (w,u)(w,u)5. The lower-bound linear program has a closed-form solution: (w,u)(w,u)6 and the upper bound is obtained analogously with

(w,u)(w,u)7

The paper proves that these bounds are sharp, that the corresponding estimators converge to the sharp partial-identification bounds, and that the rate of convergence is minimax optimal up to log factors (Khan et al., 2023).

5. Words, codes, and extremal non-overlap inequalities

In combinatorics on words, non-overlap bounds are sharp inequalities on the densities of sets that avoid suffix–prefix matches. For (w,u)(w,u)8, define (w,u)(w,u)9 as the set of words that do not overlap with any word in $1$00. If $1$01 and $1$02 are non-overlapping and $1$03, $1$04, Theorem 1.1 gives

$1$05

hence

$1$06

This inequality is sharp up to a factor of $1$07. Its proof proceeds through shift operators, a layered decomposition $1$08, a convolution inequality

$1$09

and a random-walk interpretation that converts the recursion into a summable bound (Zakharov, 23 Feb 2026).

String algorithms turn the same combinatorics into query-complexity bounds. For a text $1$10 and pattern $1$11, the non-overlapping indexing problem asks for a maximal set of non-overlapping occurrences of $1$12 in $1$13, while the range version restricts occurrences to $1$14. The paper gives an $1$15-space index with query time

$1$16

for the global problem, and an $1$17-space structure with query time

$1$18

for the range problem (0909.4893).

Coding theory generalizes these extremal problems. For $1$19-ary block codes, a $1$20-overlap-free code forbids prefix–suffix overlaps of every length $1$21. For $1$22-overlap-free codes one has

$1$23

and for $1$24-overlap-free codes with $1$25,

$1$26

Binary constructions include the Doubling Construction, the m-minimum Construction, and the Zero Block Construction, where the latter yields

$1$27

with $1$28 the $1$29-step Fibonacci sequence. In the non-overlapping case this leads to explicit lower bounds

$1$30

and to the asymptotic Levenshtein bound

$1$31

A later extension gives the general upper bound

$1$32

and, for $1$33-overlap-free codes with $1$34,

$1$35

the number of primitive words of length $1$36. That paper also constructs $1$37-overlap-free codes and codes that are simultaneously $1$38- and $1$39-overlap-free, and completes the characterization of non-expandable $1$40-overlap-free codes (Blackburn et al., 2022, Stanovnik, 2024).

6. Other manifestations and recurring techniques

Several further areas produce non-overlap bounds of the same formal type. In additive combinatorics, if $1$41 contains no non-trivial skew corner $1$42 with $1$43, then

$1$44

This Behrend-shape upper bound is proved by a two-dimensional Kelley–Meka density-increment method using vertical-segment norms, dependent random choice, almost-periodicity, and Bohr sets (Milićević, 2024).

In geometric optimization, the relaxed packing quality

$1$45

interpolates between strict non-overlap and controlled overlap. For the diagonal-distortion family $1$46, distance-based overlap yields a robust optimum at $1$47 for packing and $1$48 for covering, independent of $1$49. For volume-based overlap, the hexagonal lattice remains optimal in dimension $1$50, while in dimension $1$51 the paper gives numerical evidence for an FCC-to-BCC transition as $1$52 increases (Iglesias-Ham et al., 2014).

In constraint programming, the no-overlap constraint is NP-complete to enforce at bound consistency. The paper constructs the first bound-consistent algorithm by scanning an exact no-overlap multi-valued decision diagram (MDD): for job $1$53,

$1$54

Relaxed MDDs with bounded width then provide a polynomial-time approximation that is stronger than previously proposed precedence-detection filtering and complementary to classical propagation methods (Guichard et al., 21 Jan 2026).

In neural quantum states, overlap estimation also becomes a non-overlap-bounding problem. For normalized states, if $1$55 is the Monte Carlo estimator of $1$56 from $1$57 samples and $1$58 is the fidelity, then

$1$59

and Chebyshev’s inequality yields

$1$60

For general unnormalized states, the fidelity estimator $1$61 satisfies

$1$62

These formulas make near-orthogonality certifiable: a small estimated fidelity plus a small error bound yields a finite-sample upper bound on true overlap (Szołdra, 2023).

Taken together, these results show that non-overlap bounds are not a single theorem but a family of domain-specific inequalities with a common mathematical function. They convert absence of support, absence of shared prefixes and suffixes, or absence of simultaneous feasibility into quantitative limits on discrepancy, density, prediction accuracy, search space, or estimation error. Across the literature, the characteristic mechanisms are likelihood-ratio control, decomposition into overlap and non-overlap regions, worst-case bounding under bounded outcomes, linear and mixed-integer programming, isoperimetric and cyclic-counting arguments, spline or Lipschitz extrapolation, and diagrammatic feasibility filtering.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Non-Overlap Bounds.