---
title: Intersecting Diversity in Combinatorics & Social Metrics
url: https://www.emergentmind.com/topics/intersecting-diversity
type: topic
---

# Intersecting Diversity in Combinatorics & Social Metrics

In extremal combinatorics, **intersecting diversity** denotes a family of parameters that quantify how far an intersecting family is from a star, that is, from the most concentrated Erdős–Ko–Rado-type extremal configuration. In a distinct quantitative social-science usage, **intersecting diversity** is a normalized probability that two sampled group members differ in at least one trait when identities are treated jointly across several axes. Both usages replace one-dimensional size counts by measures of spread, concentration, or aggregate identity composition, but they do so in different mathematical settings and with different extremal or probabilistic objectives [1811.01111] [2509.14237].

## 1. Core combinatorial meaning

Let \([n]=\{1,\dots,n\}\), and let \(\mathcal F\subseteq \binom{[n]}{k}\). The family \(\mathcal F\) is **intersecting** if
\[
F\cap F'\neq \emptyset \qquad \text{for all }F,F'\in \mathcal F.
\]
A **full star** is
\[
\mathcal S_x=\{F\in \tbinom{[n]}{k}: x\in F\},
\]
and any subfamily of a full star is a star. The degree of an element \(i\in[n]\) is the number of members of \(\mathcal F\) containing \(i\), and the maximum degree is
\[
\Delta(\mathcal F)=\max_{i\in[n]} d_i.
\]
The standard **diversity** is
\[
\gamma(\mathcal F)=|\mathcal F|-\Delta(\mathcal F),
\]
equivalently,
\[
\gamma(\mathcal F)=\min_{x\in[n]}\bigl|\{F\in\mathcal F:x\notin F\}\bigr|.
\]
Thus diversity is exactly the number of sets avoiding a most popular element; it is \(0\) precisely for stars, so it measures how much of the family lies outside a largest star [1709.02829] [1811.01111].

This parameter is natural because the Erdős–Ko–Rado theorem identifies stars as the largest intersecting \(k\)-uniform families. Diversity asks a different extremal question: not how large an intersecting family can be, but how large its non-star residue can be. In shifted families, the degrees satisfy \(d_{\mathcal F}(1)\ge d_{\mathcal F}(2)\ge\cdots\), and then \(\gamma(\mathcal F)\) is simply the number of sets avoiding \(1\); this makes diversity a stability parameter as well as an extremal one [1811.01111].

## 2. Ordinary diversity and its extremal constructions

A central benchmark is Frankl’s conjectured bound
\[
\gamma(\mathcal F)\le \binom{n-3}{k-2},
\]
motivated by the classical “two out of three” family
\[
\mathcal A_2=\Bigl\{F\in \binom{[n]}{k}: |F\cap [3]|\ge 2\Bigr\},
\]
for which
\[
\gamma(\mathcal A_2)=\binom{n-3}{k-2}.
\]
The same binomial scale also appears in the “pure triangle” family
\[
\mathcal F_{123}=\Bigl\{F\in\binom{[n]}{k}: |F\cap\{1,2,3\}|=2\Bigr\},
\]
which has ordinary diversity exactly \(\binom{n-3}{k-2}\) [1709.02829] [2308.14028].

The asymptotic large-\(n\) theory was developed in stages. A 2017 result proved that there exists an absolute constant \(C\) such that for \(n>Ck\), every intersecting \(\mathcal F\subset \binom{[n]}{k}\) satisfies
\[
\gamma(\mathcal F)\le \binom{n-3}{k-2},
\]
and if equality holds then \(\mathcal F\) is a subfamily of an isomorphic copy of \(\mathcal A_2\) [1709.02829]. A later improvement showed that if \(n>36k\), then every intersecting family satisfies
\[
\gamma(\mathcal F)<\binom{n-3}{k-2},
\]
improving the previous best threshold \(n>72k\); the proof proceeds through strong lower bounds on \(\Delta(\mathcal F)\) for large intersecting families [2304.11089].

The conjectured threshold \(n>3k\), however, is not valid in full generality. For sufficiently large \(k\) and
\[
3k<n<\left(2+\frac{1}{13}\right)k,
\]
there exists an intersecting family \(\mathcal F\subset\binom{[n]}{k}\) such that
\[
\operatorname{div}(\mathcal F) > \binom{n-3}{k-2},
\]
which disproves Frankl’s \(n>3k\) threshold and also a stronger conjecture of Kupavskii in the \(r=1\) case [1804.11269]. In the nonuniform setting \(n=2k+1\), Huang’s conjecture was likewise disproved by explicit intersecting families \(\mathcal P_k\) and \(\mathcal R_k\) whose diversity exceeds
\[
\sum_{i=k+1}^{2k}\binom{2k}{i},
\]
including the clean identity
\[
\operatorname{div}(\mathcal P_k)=\sum_{i=k+1}^{2k}\binom{2k}{i}+1
\]
[1903.03585].

A further development is strong stability. If
\[
\gamma(\mathcal F)=(1-a)\binom{n-3}{k-2},\qquad 0<a<1,
\]
then, under the hypotheses of the stability theorem in [2308.14028], there exists a triple \(\{u,v,w\}\subset[n]\) such that \(\mathcal F\) differs from a pure triangle family \(\mathcal F_{uvw}\) by quantitatively controlled error terms. This makes the triangle configuration the local model for near-extremal ordinary diversity [2308.14028].

## 3. Weighted, higher-order, and operational variants

Several generalizations retain the same star-versus-spread interpretation while altering the objective.

| Variant | Definition | Representative result |
|---|---|---|
| Ordinary diversity | \(\gamma(\mathcal F)=|\mathcal F|-\Delta(\mathcal F)\) | For large \(n\), the benchmark scale is \(\binom{n-3}{k-2}\) [1709.02829] |
| \(C\)-weighted diversity | \(d_C(\mathcal F)=|\mathcal F|-C\Delta(\mathcal F)\) | For \(1<C<\tfrac32\) and \(n\ge \frac{42}{3-2C}k\), \(\mathcal F_{123}\) uniquely maximizes \(\gamma_C\) [2308.14028] |
| Double-diversity | \(\gamma_2(\mathcal F)=\min_{x,y}|\mathcal F(\bar x,\bar y)|\) | For \(n\ge 13k^2\), the Fano \(k\)-graph is the unique extremal family [2212.11650] |
| Symmetric-difference spread | \(\mathcal{SD}(\mathcal F)=\{F\triangle G:F,G\in\mathcal F\}\) | For \(n\ge 60k^{3/2}\), \(k\ge 50\), stars maximize \(|\mathcal{SD}(\mathcal F)|\) [2606.20043] |

The \(C\)-weighted theory interpolates between size and diversity. When \(C=0\), maximizing \(d_C\) is just the Erdős–Ko–Rado problem. When \(C=1\), it becomes ordinary diversity. For larger \(C\), the penalty on high degree becomes stronger. One paper determines the maximal families for \(C\in[0,\frac73)\) for large \(n\): stars are optimal for \(0\le C<1\); “two out of three” configurations govern the interval \(1\le C<\frac32\); and Fano-plane-based families \(A_{\mathbb F}\) and \(A_{\mathbb F^+}\) govern \(\frac32\le C<\frac73\). The same work records the corresponding growth-rate drops
\[
\Theta(n^{r-1}),\qquad \Theta(n^{r-2}),\qquad \Theta(n^{r-3})
\]
across these phases [2306.00384].

Higher-order diversity replaces deletion of one vertex by deletion of several. For \(l\)-diversity,
\[
\gamma_l(\mathcal F)=\min_{S\in\binom{[n]}{l}}|\mathcal F(S)|,
\]
where \(\gamma_1\) is ordinary diversity and \(\gamma_2\) is double-diversity. The exact double-diversity theorem states that if \(n\ge 13k^2\), then
\[
\gamma_2(\mathcal F)\le 2\binom{n-5}{k-3}-\binom{n-7}{k-5},
\]
with equality if and only if \(\mathcal F\) is isomorphic to the Fano \(k\)-graph. For triple diversity, the best stated result is an upper bound
\[
\gamma_3(\mathcal F)\le 3\binom{n-7}{k-4}+108k\binom{n-8}{k-5}\qquad (n>71k^2),
\]
and for \(l\ge4\) the theory is much coarser [2212.11650].

A related operational statistic is the symmetric-difference family \(\mathcal{SD}(\mathcal F)\). Although it is not itself a diversity parameter, it is presented as a symmetric-difference analogue of classical extremal questions for intersecting families and is stated to be closely related to intersecting diversity. In the proved range \(n\ge 60k^{3/2}\), \(k\ge 50\),
\[
|\mathcal{SD}(\mathcal F)|\le \sum_{\ell=0}^{k-1}\binom{n-1}{2\ell},
\]
with equality only for stars [2606.20043].

## 4. Structural methods and extensions beyond set systems

The theory is driven by structural reduction. One influential large-\(n\) approach uses the Dinur–Friedgut junta method: a sufficiently large intersecting family is essentially contained in a bounded-size junta, and Proposition 4 in [1709.02829] gives a dichotomy for intersecting juntas that either places the family inside an \(\mathcal A_2\)-type structure or forces diversity to be smaller than the benchmark. The exact extremal comparison is then completed by a cross-intersecting lemma proved with Kruskal–Katona and lexicographic compression [1709.02829].

A different line uses **shifting ad extremis**. The 2023 improvement from \(72k\) to \(36k\) relies on shifting until only a small set of resistant pairs remains, followed by fiber-counting and cross-intersecting inequalities. A key structural claim is that the graph of shift-resistant pairs cannot contain three pairwise disjoint edges. This converts size information into a lower bound on \(\Delta(\mathcal F)\), and then into an upper bound on diversity [2304.11089].

For \(C\)-weighted diversity, a variant of Frankl’s Delta-system method called the **flower base** plays the same role. A flower with threshold \(\alpha\) is a family with a common core \(Y\) whose petals outside \(Y\) have transversal number \(>\alpha\), and the Flower Lemma states that every sufficiently large family of \(r\)-sets contains such a flower. The flower base \(\mathfrak B F\) is Sperner, intersecting when \(F\) is intersecting, covers every edge of \(F\), and has bounded size. This makes it possible to classify the small skeleton rather than the original family, producing the star / two-out-of-three / Fano-plane trichotomy [2306.00384].

Cross-intersecting families admit a parallel diversity theory. In that setting, one writes the degree part and diversity part of each family as \(\mathcal A_\Delta,\mathcal A_\gamma,\mathcal B_\Delta,\mathcal B_\gamma\). A 2026 structural theorem extends Kupavskii’s theorem from a single intersecting family to large cross-intersecting pairs and shows that, once the diversity parts are fixed, the maximal degree parts are the maximal cross-intersecting extensions. A technical innovation is the \(S_{U,V}^Q\)-shift, designed to preserve both cross-intersection and local structural constraints [2606.20085].

The same philosophy has been exported to other combinatorial categories. For \(k\)-subspaces of \(V(n,q)\), diversity is the number of subspaces not containing the most popular \(1\)-dimensional subspace. In wide parameter regimes, the family \(G_2\) has the largest diversity, with
\[
\gamma(G_2)=q^k\binom{n-3}{k-2}_q,
\]
and the same paper proves a Frankl-type degree-diversity theorem with extremal families \(G_i\) [2605.02698]. A complementary 2026 paper studies ordered point-degrees \(d_1(\mathcal F)\ge d_2(\mathcal F)\ge\cdots\) and proves, for intersecting families of \(k\)-subspaces with \(n\ge 2k+1\),
\[
d_{\points{k}^2}(\mathcal F)\le \qbinom{n-2}{k-2},
\]
together with a corrected \(q\)-analogue of the Huang–Rao \((k+2)\)-th degree theorem for fixed \(q\), sufficiently large \(k\), and \(n>3k\) [2606.08709].

For permutation families \(\mathcal F\subset \mathcal S_n\), intersecting means that every pair of permutations agrees in at least one position. Diversity is then the minimum number of permutations whose deletion results in a star. For \(n\ge 500\),
\[
\gamma(\mathcal F)\le (n-3)(n-3)!,
\]
and this is sharp, with equality attained by triangle families. The proof uses the spread approximation method of Kupavskii and Zakharov together with Füredi’s pseudo-sunflower theorem [2501.06731].

## 5. A distinct probabilistic metric for intersecting demographic traits

A separate literature uses **intersecting diversity** to quantify aggregate identity composition in groups. Consider a group \(G\) of \(N\) individuals and \(T\) traits. If the possible aggregate identities are
\[
c=(c_1,\ldots,c_T)\in \boldsymbol{\mathcal C},\qquad C:=|\boldsymbol{\mathcal C}|=v_1v_2\cdots v_T,
\]
and \(p_c\) is the proportion of the group with identity \(c\), then intersecting diversity is defined by
\[
\mathcal D:=\frac{C}{C-1}\left(1-\sum_{c\in\boldsymbol{\mathcal C}}p_c^2\right)
=\frac{C}{C-1}\bigl(1-P(X=T)\bigr),
\]
where \(X\) is the number of shared traits between two independently sampled individuals. Equivalently, \(\mathcal D\) is a normalized probability that two sampled individuals differ in at least one trait [2509.14237].

The companion metric is **shared identity**,
\[
\mathcal S:=\frac{1}{T}\sum_{t=1}^T \sum_{v=1}^{v_t}
\left(\sum_{\substack{c\in\boldsymbol{\mathcal C}\\ c_t=v}} p_c\right)^2
=\frac{E(X)}{T}.
\]
The paper also defines \(\mathcal S_N\) using sampling without replacement and proves
\[
\mathcal S=\left(1-\frac1N\right)\mathcal S_N+\frac1N.
\]
Its main theorem establishes the bounds
\[
1-\frac{C-1}{C}\mathcal D\le \mathcal S\le 1-\frac{C-1}{TC}\mathcal D,
\]
the lower bound
\[
\mathcal S_{\min}=\frac1T\sum_{t=1}^T \frac1{v_t}\le \mathcal S,
\]
and the gradient inequality
\[
\langle \nabla \mathcal D,\nabla \mathcal S\rangle
\le -\frac{4C}{C-1}\mathcal S_{\min}<0.
\]
These formulas show that intersecting diversity and shared identity are structurally anti-correlated and that there is no clear “optimal” point maximizing both metrics simultaneously [2509.14237].

The paper works out three case studies. In Hollywood-movie crews and Survivor tribes, the empirical \((\mathcal D,\mathcal S)\) clouds lie inside the admissible region and display the predicted anti-correlation. In Survivor, seasons beginning in 2020 are reported as more diverse overall than earlier seasons, while trait-by-trait analysis shows that race/ethnicity diversity increased and age diversity decreased. For North American companies, pairwise comparisons in which one company had both higher \(\mathcal D\) and higher \(\mathcal S\) yielded a higher industry-adjusted EBIT margin only \(37\%\) of the time, and the paper states that this points in the opposite direction of the usual “more diversity + more shared identity = better performance” narrative [2509.14237].

## 6. Broader intersectional and networked interpretations

Several adjacent literatures study intersecting forms of diversity without using the extremal-set-theoretic parameter. Information decomposition gives one such formulation. Using partial information decomposition, one paper interprets **intersectional synergy** as information about an outcome available only from the combination of identities, not from any identity alone. In U.S. census microdata, the relationship race + sex \(\to\) income is reported as about \(51\%\) synergy, about \(42\%\) redundancy, and about \(7\%\) unique information total; the same paper uses synthetic data to show that linear regression with multiplicative interaction coefficients does not distinguish genuinely synergistic effects from redundant ones [2106.10338].

In HCI, type abstraction is used to avoid the combinatorial explosion of intersectional analysis. A formal compositional theorem states
\[
\iMag[D\Join D',State]=\iMag[D,State]\cup \iMag[D',State],
\]
so that separate analyses along different diversity dimensions can be joined by union. The claim is not that empirical work disappears, but that prior analytical artifacts can be reused when moving from one-dimensional to intersectional populations [2201.10643].

Networked collective-learning models add a further layer. In one such model, diversity is implemented as heterogeneity in payoff functions. The main result is conditional: for simple tasks, diversity consistently impairs performance, whereas for complex tasks the effect depends on network density—diversity hurts in sparse networks and helps in dense networks, including the fully connected limit [2306.17812]. A related empirical study of everyday geography uses 49 mobility surveys, 385,000 respondents, and 1,711,000 trips to examine hourly intersectional urban patterns across gender, age, and education; it reports that in strongly daytime-attractive and strongly nighttime-decreasing districts, dominant groups are much more synchronous than non-dominant groups [2106.15492].

A broader socio-technical synthesis appears in work on the “Internet of Us,” which models each participant by a vector of visible and invisible profile features and distinguishes **descriptive diversity** from **prescriptive diversity**. In that framework, diversity-aware AI is used to diversify rankings, mediate norms, and allow communities to choose which profile characteristics matter for diversification in their setting. This suggests a pragmatic, system-design interpretation of intersecting diversity as the joint management of multiple dimensions of difference rather than the optimization of a single scalar score [2503.16448].

Across these literatures, a common pattern persists. Intersecting diversity is not a synonym for size, heterogeneity, or fairness in the abstract. In extremal set theory it is a calibrated distance from a star; in the aggregate-identity metric it is a normalized probability of pairwise difference; in adjacent intersectional and networked work it becomes a way to formalize irreducible joint effects, compositional analysis, or context-dependent benefits of heterogeneous groups. The shared theme is that one-dimensional summaries are insufficient once intersection, concentration, or joint identity structure becomes the main object of study.

Source: https://www.emergentmind.com/topics/intersecting-diversity