Non-Overlap Bounds in Theory & Applications
- Non-Overlap Bounds are quantitative statements that limit discrepancies when overlap between entities is absent, with key applications in causal inference, combinatorics, and coding theory.
- They employ methodologies like uniform likelihood-ratio bounds, linear programming, and Lipschitz extrapolation to derive sharp, domain-specific inequalities and partial-identification intervals.
- These bounds offer practical tools for error estimation, sensitivity analysis, and feasibility constraints across disciplines, ensuring rigorous assessment of non-overlap scenarios.
Non-overlap bounds are quantitative statements that limit what can occur when overlap, positivity, or compatibility between objects is absent or explicitly restricted. Across the literatures represented here, they appear as partial-identification intervals for causal effects, information-theoretic bounds on discrepancy between treated and control populations, sharp extremal inequalities for words and codes, geometric density limits under controlled intersection, feasibility bounds for scheduling constraints, and Monte Carlo error bounds for near-orthogonal quantum states (D'Amour et al., 2017, Susmann et al., 24 Sep 2025, Zakharov, 23 Feb 2026).
1. Formal notions of overlap and non-overlap
In observational causal inference, overlap is expressed through the propensity score. With binary treatment, population overlap requires with probability $1$, while strict overlap strengthens this to
for some fixed . In off-policy evaluation, the analogous support condition is , and the no-overlap region is . In scheduling, the no-overlap constraint on a single machine is the disjunction for every pair of jobs (D'Amour et al., 2017, Khan et al., 2023, Guichard et al., 21 Jan 2026).
In combinatorics on words, overlap is defined through prefix–suffix coincidence. For two words , overlaps if there exists $1$0 such that the suffix of $1$1 of length $1$2 equals the prefix of $1$3 of length $1$4. In coding theory, $1$5 and $1$6 have a $1$7-overlap if the prefix of $1$8 of length $1$9 is identical to the suffix of 0 of length 1, and a 2-overlap-free code forbids such overlaps for every 3 with 4. These definitions specialize to classical non-overlapping codes when all proper overlaps are forbidden (Zakharov, 23 Feb 2026, Blackburn et al., 2022, Stanovnik, 2024).
In geometric and physical settings, overlap is measured differently but plays the same structural role. In the sphere-packing generalization with limited overlap, one controls either a distance-based overlap
5
or a volume-based overlap
6
For neural quantum states, the overlap is the Hilbert-space inner product 7, and non-overlap corresponds to orthogonality or near-orthogonality (Iglesias-Ham et al., 2014, Szołdra, 2023).
2. Overlap as a global restriction in high-dimensional causal inference
A central modern use of non-overlap bounds is to reinterpret strict overlap as a bound on how different treated and control covariate distributions can be. If 8 and 9 are the conditional covariate distributions under treatment and control, then strict overlap is equivalent to a uniform likelihood-ratio bound
0
with
1
This implies bounded 2-divergences, including bounded 3-divergence and bounded Kullback–Leibler divergence between 4 and 5. When 6,
7
and
8
These are global non-overlap bounds: if strict overlap holds, treated and control covariate distributions cannot be too far apart in KL or 9 divergence (D'Amour et al., 2017).
The same framework yields explicit mean-imbalance bounds. If 0 and 1 denote covariate means and 2 the covariance matrices, then strict overlap implies
3
Consequently,
4
If the smaller operator norm is 5, the average absolute mean difference goes to zero as 6. The paper also shows that strict overlap bounds treatment-prediction accuracy: 7 for every classifier 8. In this sense, strict overlap excludes asymptotically perfect classification and forces the average per-covariate discriminatory information to vanish in high dimension.
3. Partial identification, sensitivity analysis, and structural non-overlap
When overlap fails, one line of work retains the average treatment effect and replaces point identification by partial identification. With bounded outcomes 9, the average treatment effect can be decomposed into an identifiable trimmed effect
0
and a worst-case contribution from the non-overlap sets 1 and 2. This yields the non-overlap bounds
3
so that 4. The width is
5
namely the probability mass of the non-overlap population. Because these bounds are non-smooth, the paper introduces smooth approximations 6 and 7, derives efficient influence functions, proposes a Targeted Minimum Loss-Based estimator, and constructs a multiplier-bootstrap confidence set that is uniformly valid across all trimming and smoothing parameters (Susmann et al., 24 Sep 2025).
A related finite-sample perspective decomposes the full-sample average treatment effect as 8, where 9 is the contribution from the overlap region 0 and 1 is the contribution from the limited-overlap region 2, with 3. Under the Lipschitz class
4
the paper builds minimax confidence intervals 5 for 6. The worst-case bias and variance are controlled through the modulus of continuity 7, and a combined interval
8
is formed by adding an asymptotic interval for 9 to the minimax interval for 0. This converts limited overlap into an explicit sensitivity analysis over the smoothness constant 1 (Ma et al., 27 Nov 2025).
Another strand allows simultaneous violations of unconfoundedness and overlap. Under the robust marginal sensitivity model,
2
only a fraction 3 of units in group 4 satisfy the odds-ratio bound; the remainder may have arbitrary confounding and 5 or 6. The target estimand is the overlap-weighted average treatment effect
7
which remains identifiable because units with 8 or 9 receive zero weight. Worst-case lower and upper bounds are computed by linear programming or mixed-integer linear programming under robust sensitivity and covariate-balance constraints, and inference is based on an augmented percentile bootstrap (Cui et al., 16 Sep 2025).
Two further contributions situate non-overlap within explicit modeling architectures. In the Bayesian BART+SPL framework, the sample is split into a region of overlap 0 and a region of non-overlap 1 according to the estimated propensity score; Bayesian Additive Regression Trees are used in 2, while a spline model extrapolates into 3 with variance inflation
4
where 5 is the distance into non-overlap and 6 is the range of causal effects in 7 (Nethery et al., 2018). In fusion problems with randomized and observational data, structural non-overlap induces a feasibility gap
8
and under conditional structural non-overlap one obtains
9
Representation learning can reduce this gap by recovering overlap in a latent space, but the resulting oracle inequality still decomposes excess risk into overlap recovery, moment violation, and statistical terms (Du et al., 26 Feb 2026).
4. Off-policy evaluation beyond overlap
In off-policy evaluation, the value of a target policy is
0
and standard importance-weighting methods require the support condition 1. Without this condition, the paper decomposes the value into an identified overlap part and an unidentified no-overlap part,
2
then partially identifies 3 under the Lipschitz class
4
possibly intersected with boundedness 5. The lower-bound linear program has a closed-form solution: 6 and the upper bound is obtained analogously with
7
The paper proves that these bounds are sharp, that the corresponding estimators converge to the sharp partial-identification bounds, and that the rate of convergence is minimax optimal up to log factors (Khan et al., 2023).
5. Words, codes, and extremal non-overlap inequalities
In combinatorics on words, non-overlap bounds are sharp inequalities on the densities of sets that avoid suffix–prefix matches. For 8, define 9 as the set of words that do not overlap with any word in $1$00. If $1$01 and $1$02 are non-overlapping and $1$03, $1$04, Theorem 1.1 gives
$1$05
hence
$1$06
This inequality is sharp up to a factor of $1$07. Its proof proceeds through shift operators, a layered decomposition $1$08, a convolution inequality
$1$09
and a random-walk interpretation that converts the recursion into a summable bound (Zakharov, 23 Feb 2026).
String algorithms turn the same combinatorics into query-complexity bounds. For a text $1$10 and pattern $1$11, the non-overlapping indexing problem asks for a maximal set of non-overlapping occurrences of $1$12 in $1$13, while the range version restricts occurrences to $1$14. The paper gives an $1$15-space index with query time
$1$16
for the global problem, and an $1$17-space structure with query time
$1$18
for the range problem (0909.4893).
Coding theory generalizes these extremal problems. For $1$19-ary block codes, a $1$20-overlap-free code forbids prefix–suffix overlaps of every length $1$21. For $1$22-overlap-free codes one has
$1$23
and for $1$24-overlap-free codes with $1$25,
$1$26
Binary constructions include the Doubling Construction, the m-minimum Construction, and the Zero Block Construction, where the latter yields
$1$27
with $1$28 the $1$29-step Fibonacci sequence. In the non-overlapping case this leads to explicit lower bounds
$1$30
and to the asymptotic Levenshtein bound
$1$31
A later extension gives the general upper bound
$1$32
and, for $1$33-overlap-free codes with $1$34,
$1$35
the number of primitive words of length $1$36. That paper also constructs $1$37-overlap-free codes and codes that are simultaneously $1$38- and $1$39-overlap-free, and completes the characterization of non-expandable $1$40-overlap-free codes (Blackburn et al., 2022, Stanovnik, 2024).
6. Other manifestations and recurring techniques
Several further areas produce non-overlap bounds of the same formal type. In additive combinatorics, if $1$41 contains no non-trivial skew corner $1$42 with $1$43, then
$1$44
This Behrend-shape upper bound is proved by a two-dimensional Kelley–Meka density-increment method using vertical-segment norms, dependent random choice, almost-periodicity, and Bohr sets (Milićević, 2024).
In geometric optimization, the relaxed packing quality
$1$45
interpolates between strict non-overlap and controlled overlap. For the diagonal-distortion family $1$46, distance-based overlap yields a robust optimum at $1$47 for packing and $1$48 for covering, independent of $1$49. For volume-based overlap, the hexagonal lattice remains optimal in dimension $1$50, while in dimension $1$51 the paper gives numerical evidence for an FCC-to-BCC transition as $1$52 increases (Iglesias-Ham et al., 2014).
In constraint programming, the no-overlap constraint is NP-complete to enforce at bound consistency. The paper constructs the first bound-consistent algorithm by scanning an exact no-overlap multi-valued decision diagram (MDD): for job $1$53,
$1$54
Relaxed MDDs with bounded width then provide a polynomial-time approximation that is stronger than previously proposed precedence-detection filtering and complementary to classical propagation methods (Guichard et al., 21 Jan 2026).
In neural quantum states, overlap estimation also becomes a non-overlap-bounding problem. For normalized states, if $1$55 is the Monte Carlo estimator of $1$56 from $1$57 samples and $1$58 is the fidelity, then
$1$59
and Chebyshev’s inequality yields
$1$60
For general unnormalized states, the fidelity estimator $1$61 satisfies
$1$62
These formulas make near-orthogonality certifiable: a small estimated fidelity plus a small error bound yields a finite-sample upper bound on true overlap (Szołdra, 2023).
Taken together, these results show that non-overlap bounds are not a single theorem but a family of domain-specific inequalities with a common mathematical function. They convert absence of support, absence of shared prefixes and suffixes, or absence of simultaneous feasibility into quantitative limits on discrepancy, density, prediction accuracy, search space, or estimation error. Across the literature, the characteristic mechanisms are likelihood-ratio control, decomposition into overlap and non-overlap regions, worst-case bounding under bounded outcomes, linear and mixed-integer programming, isoperimetric and cyclic-counting arguments, spline or Lipschitz extrapolation, and diagrammatic feasibility filtering.