Papers
Topics
Authors
Recent
Search
2000 character limit reached

Quantitative Intersectional Data (QUINTA)

Updated 12 July 2026
  • QUINTA is a set of quantitative frameworks that jointly model intersecting social identities to measure and govern multiple protected attributes.
  • It employs information-theoretic metrics like mutual information to price informational harm and assess data leakage in a model-agnostic manner.
  • QUINTA supports fairness auditing in machine learning by constructing joint distributions over protected attributes to identify and mitigate intersectional biases.

Searching arXiv for the cited QUINTA-related papers and closely related intersectional methodology papers. arxiv_search: (Garrido-Merchán, 19 Jul 2025) arxiv_search: (Garrido-Merchán, 13 Jan 2026) arxiv_search: (Souligne et al., 6 Apr 2026) arxiv_search: (Morina et al., 2019) arxiv_search: (Foulds et al., 2018) arxiv_search: (Varley et al., 2021) Quantitative Intersectional Data (QUINTA) denotes a family of quantitative frameworks for representing, measuring, and governing data about intersecting social identities. In recent work, the acronym is expanded in several ways—“QUantitative INTersectional data lAbeling and valuation,” “QUANTitative INtersectionAl data,” “QUantitative INtersectional TAxonomy,” and “QUANTitative InTersectional Analysis”—but the common premise is consistent: protected attributes such as race, gender, disability, class, sexuality, age, education, or political identity are modeled jointly, and the relevant quantities are defined on the joint distribution or on explicitly constructed intersectional groups rather than on single axes in isolation (Garrido-Merchán, 19 Jul 2025, Garrido-Merchán, 13 Jan 2026, Souligne et al., 6 Apr 2026, Bogucka et al., 27 Apr 2026).

1. Terminological range and analytical object

Across the literature, QUINTA is not a single algorithm. It is a recurrent label for quantitative intersectional workflows in data valuation, fairness auditing, survey analysis, historical NLP, educational measurement, incident analysis, and reflexive AI research. The shared analytical move is to define an intersectional object explicitly: a joint protected profile such as S=(S1,,Sk)S=(S_1,\dots,S_k) or I=(I1,,Im)I=(I_1,\dots,I_m), a Cartesian product of group axes G=A1××ApG=A_1\times\cdots\times A_p, or a finite set of aggregate identities cCc\in\mathcal C with proportions pcp_c (Garrido-Merchán, 19 Jul 2025, Garrido-Merchán, 13 Jan 2026, Hoogstra et al., 11 Aug 2025).

Use of QUINTA Primary object Representative paper
Data valuation Joint protected profile SS or II (Garrido-Merchán, 19 Jul 2025, Garrido-Merchán, 13 Jan 2026)
Fairness auditing Intersectional groups from multiple protected attributes (Souligne et al., 6 Apr 2026, Morina et al., 2019)
Underrepresentation correction Sensitive groups G1,,GkG_1,\dots,G_k and their intersections (Diana et al., 2023)
Empirical social analysis Joint strata over gender, age, education, race, sexuality, or class (Vallée et al., 2021, Cohen et al., 3 Mar 2026)
Reflexive AI/DS practice Tuple (P,D,M,δ)(P,D,M,\delta) over positionality, data, methods, and reflexive metadata (Boyd, 19 Sep 2025)

This terminological variation matters because the literature uses “QUINTA” both narrowly and broadly. In some papers it names a formal pricing rule for the social cost of intersectional inference; in others it names an auditing toolkit, a survey-analysis protocol, or a reflexive methodological paradigm. Taken together, these uses suggest that QUINTA is best understood as a quantitative intersectional methodology class rather than as a single standardized framework (Boyd, 19 Sep 2025).

2. Information-theoretic valuation and Pigouvian pricing

A central QUINTA line of work formulates intersectional privacy as an externality. The core claim is that the “social harm” of releasing a datum XX derives exclusively from how much it reduces uncertainty about a full protected profile I=(I1,,Im)I=(I_1,\dots,I_m)0 taken jointly. The corresponding metric is Shannon mutual information,

I=(I1,,Im)I=(I_1,\dots,I_m)1

which quantifies the expected drop in entropy about the protected profile when I=(I1,,Im)I=(I_1,\dots,I_m)2 is observed (Garrido-Merchán, 19 Jul 2025).

Under the axioms Intersectional Externality, No-Harm/No-Cost, Monotonicity, and Additivity, the price of a datum is forced into the linear form

I=(I1,,Im)I=(I_1,\dots,I_m)3

with pure surcharge

I=(I1,,Im)I=(I_1,\dots,I_m)4

where I=(I1,,Im)I=(I_1,\dots,I_m)5 is the data-collection cost and I=(I1,,Im)I=(I_1,\dots,I_m)6 is the regulator’s social-cost parameter per unit of information. A related pointwise formulation prices each observed data point by

I=(I1,,Im)I=(I_1,\dots,I_m)7

thereby translating informational harm into dollars per nat or per bit (Garrido-Merchán, 19 Jul 2025, Garrido-Merchán, 13 Jan 2026).

These formulations are explicitly model-agnostic. The surcharge depends on I=(I1,,Im)I=(I_1,\dots,I_m)8 or on the posterior/prior pair I=(I1,,Im)I=(I_1,\dots,I_m)9 and G=A1××ApG=A_1\times\cdots\times A_p0, not on any specific downstream predictive model. The same rule can therefore be implemented whether the joint distribution is estimated via histograms, kernel-density estimation, k-nearest-neighbour estimators, logistic regression, random forests, auto-encoders, GANs, mixture models, or variational bounds such as MINE (Garrido-Merchán, 19 Jul 2025, Garrido-Merchán, 13 Jan 2026).

The operational recipe is straightforward. One discretizes or otherwise models the joint intersectional space, collects an audited sample of G=A1××ApG=A_1\times\cdots\times A_p1 or G=A1××ApG=A_1\times\cdots\times A_p2, estimates the joint and conditional distributions, computes G=A1××ApG=A_1\times\cdots\times A_p3, and multiplies by a policy parameter. The papers propose calibration by setting a statutory maximum penalty G=A1××ApG=A_1\times\cdots\times A_p4 for full disclosure and choosing G=A1××ApG=A_1\times\cdots\times A_p5. They also give concrete examples: for a race G=A1××ApG=A_1\times\cdots\times A_p6 gender G=A1××ApG=A_1\times\cdots\times A_p7 disability space with G=A1××ApG=A_1\times\cdots\times A_p8 joint categories, G=A1××ApG=A_1\times\cdots\times A_p9 nats under a uniform prior; a feature leaking cCc\in\mathcal C0 nats would incur a surcharge of \$c\in\mathcal C$1\lambda=100{,}000\ \$c\in\mathcal C$2, whereas a timestamp leaking $c\in\mathcal C$3 nats would cost about \$2,000 (Garrido-Merchán, 19 Jul 2025).

A further extension introduces normative weights cCc\in\mathcal C4 on intersectional groups, yielding a weighted mutual information

cCc\in\mathcal C5

with cCc\in\mathcal C6. This makes the pricing rule not merely model-agnostic but explicitly norm-sensitive: cCc\in\mathcal C7 and cCc\in\mathcal C8 become the policy knobs through which contested social values and redistributive aims are embedded into market pricing (Garrido-Merchán, 13 Jan 2026).

3. Fairness auditing, estimation, and mitigation in machine learning

A second major QUINTA strand addresses intersectional fairness in classification. The basic setup constructs the full intersectional domain as the Cartesian product of protected attributes and measures disparities on joint subgroups rather than on marginal attributes alone. Formal criteria include intersectional statistical parity, equal opportunity, false-positive parity, equalized odds, and multiplicative notions such as cCc\in\mathcal C9-differential fairness, where the worst-case log-ratio between subgroup probabilities is bounded by pcp_c0 (Morina et al., 2019).

Several papers emphasize that intersectional estimation is statistically difficult because data sparsity grows rapidly with the number of protected attributes and categories. Proposed responses include beta-binomial smoothing, bootstrap confidence intervals, Bayesian posterior estimation, and hierarchical Bayesian models that shrink sparse group estimates toward a global predictor. In the Bayesian setting, posterior samples of group-conditional probabilities induce posterior samples of fairness metrics such as differential fairness and subgroup fairness, producing uncertainty-aware estimates rather than unstable plug-in ratios (Foulds et al., 2018, Morina et al., 2019).

The clinical toolkit FairLogue operationalizes this logic in both observational and counterfactual settings. Its observational framework extends demographic parity, equalized odds, and equal opportunity difference to intersectional groups, while its generalized counterfactual framework evaluates fairness under interventions on group membership and compares observed disparities to permutation-based null distributions through a “u-value.” In a glaucoma surgery prediction task on the All of Us Controlled Tier V8 cohort, observational analysis found an intersectional demographic parity gap of pcp_c1, an equalized-odds true positive rate gap of pcp_c2, and a false positive rate gap of pcp_c3, all larger than the corresponding single-axis gaps; the counterfactual analysis, however, produced u-values near pcp_c4, suggesting that the observed disparities were consistent with chance after conditioning on covariates (Souligne et al., 6 Apr 2026).

Mitigation methods span post-processing, reweighting, and data augmentation. One post-processing framework constructs a randomized thresholding derived predictor with subgroup-specific thresholds and flip probabilities, then solves a linear program to satisfy chosen fairness constraints with minimal expected loss (Morina et al., 2019). For underrepresentation bias, another QUINTA formulation assumes a small unbiased sample and a larger biased sample in which positive examples are filtered at unknown group-wise drop-out rates pcp_c5. It estimates the pcp_c6, defines a reweighting function on positive examples, and proves agnostic PAC learnability under intersectional bias for finite-VC hypothesis classes (Diana et al., 2023).

Synthetic augmentation is handled differently in “Synthetic Data Generation for Intersectional Fairness by Leveraging Hierarchical Group Structure,” which treats intersectional groups as intersections of parent categories. The method learns a generative transformation from immediate parent groups to scarce child groups using an MMD objective and then trains downstream classifiers under equal sampling across groups. On four datasets, the reported empirical pattern is a Balanced Accuracy drop of at most pcp_c7 points versus the unconstrained baseline, worst-case FPR reductions of up to pcp_c8, reductions in differential fairness on Hate Speech from about pcp_c9 to SS0, reductions in SS1 from about SS2 to SS3 on Anxiety and from about SS4 to SS5 on Twitter Hate Speech, and no method-induced “leveling down” (Maheshwari et al., 2024).

A recurrent caution is that handling underrepresentation is not purely technical. Reweighting, resampling, GAN-based augmentation, and counterfactual augmentation all have normative implications, and one paper explicitly warns that synthetic face-image generation can reproduce disturbing tropes analogous to historical “face painting,” while unsupervised relabeling of sensitive categories can undermine self-identification and agency (Wang et al., 2022).

4. QUINTA as an empirical workflow across domains

Outside ML fairness, QUINTA appears as a general empirical protocol for intersectional measurement. In historical NLP, it combines OCR correction, tokenization, distributional semantics, and bias metrics. “Measuring Intersectional Biases in Historical Documents” uses Pointwise Mutual Information, WEAT effect sizes, and a two-dimensional mapping of descriptors into a gender SS6 race plane. The study reports that BPE yields average nearest-neighbour Jaccard similarity about SS7 but only about SS8 misspelling compatibility, whereas spaCy preprocessing yields about SS9 stability and about II0 compatibility; it also reports that white females are strongly associated with “family,” “weak,” and “appearance,” white males with “career,” “strong,” and “manual labour,” and that more than II1 of strongly “beautiful” adjectives fall into the white-female quadrant only (Borenstein et al., 2023).

In urban geography, QUINTA is used to study hourly spatial distributions across gender, age, and education. The French mobility-survey application relies on II2 origin–destination surveys, II3 respondents, II4 trips, II5 districts, and II6 hourly snapshots. It normalizes district-level group counts into probability-like temporal profiles, computes a Euclidean-based mismatch index, and applies Ward clustering. The reported results include five district temporal profiles, gender-based mismatch in II7 of districts, at least one age-based mismatch in II8, at least one education-based mismatch in II9, and a synchronization contrast in which dominant groups are more synchronous than non-dominant groups in high-flux districts but not in stable districts (Vallée et al., 2021).

In population health, the long-COVID study operationalizes intersectional strata over sex, race/ethnicity, education, and sexual and gender minority status, then combines weighted descriptive prevalence estimates with survey-corrected logistic regressions containing two-way and three-way interactions. The descriptive results reveal a “step-like pattern” in which college-educated men have the lowest prevalence of long COVID and women without college educations the highest, while the regression results show, among other effects, odds ratios of G1,,GkG_1,\dots,G_k0 for Female, G1,,GkG_1,\dots,G_k1 for BA or higher, G1,,GkG_1,\dots,G_k2 for SGM, and an interaction odds ratio of G1,,GkG_1,\dots,G_k3 for Female G1,,GkG_1,\dots,G_k4 Hispanic in the long-COVID model (Cohen et al., 3 Mar 2026).

Educational research uses a related but conceptually distinct QUINTA under the label of critical quantitative intersectionality. The physics-learning study fits hierarchical linear models with race, gender, and course-type effects, and then evaluates equity through two operationalizations: Equality of Learning and Equity of Individuality. The reported result is deliberately non-convergent: collaborative instruction improved gains for all groups by about G1,,GkG_1,\dots,G_k5 percentage points, satisfying Equity of Individuality, but the gaps relative to White men remained, so Equality of Learning was not achieved (Dusen et al., 2018).

A more recent evaluation framework for LLM empathy uses controlled prompt generation over intersectional social categories such as race, pronoun, sexuality, education, religion, age, and socioeconomic status, then measures Cognitive Empathy by five-way multiple-choice accuracy, affective empathy through VAD scores, and response appropriateness through Epitome-derived scores. In the initial evaluation sample, intersectional cell variances were all below G1,,GkG_1,\dots,G_k6, no cell mean differed significantly from the grand mean after correction, and one outlier cell—Lesbian plus “she”—showed an interpretation score about G1,,GkG_1,\dots,G_k7 standard deviations above the overall mean (Formanek et al., 2024).

5. Higher-order structure: synergy, diversity, and uncertainty

Some QUINTA frameworks move beyond groupwise disparity measurement to quantify higher-order structure directly. Partial Information Decomposition is used to separate redundant, unique, and synergistic contributions of intersecting identities to an outcome. For two predictors G1,,GkG_1,\dots,G_k8 and target G1,,GkG_1,\dots,G_k9,

(P,D,M,δ)(P,D,M,\delta)0

This formalism is used to argue that some intersectional effects are irreducible to single identities considered individually. In U.S. census-style analyses, the reported decomposition of (P,D,M,δ)(P,D,M,\delta)1 yields about (P,D,M,δ)(P,D,M,\delta)2 redundant information, (P,D,M,δ)(P,D,M,\delta)3 synergistic information, and (P,D,M,δ)(P,D,M,\delta)4 unique information, while a synthetic-data comparison shows that standard linear regression with multiplicative interaction terms fails to distinguish truly synergistic effects from redundant ones (Varley et al., 2021).

Another line defines multidimensional diversity metrics directly on aggregate identities. Intersecting Diversity,

(P,D,M,δ)(P,D,M,\delta)5

is a normalized probability that two random draws differ in at least one trait. Shared Identity,

(P,D,M,δ)(P,D,M,\delta)6

is the expected fraction of traits shared by two draws. These measures are shown to be anti-correlated, with admissible pairs constrained to a polygonal region and with no single point that maximizes both metrics simultaneously. The case studies on Hollywood films, Survivor tribes, and North American companies illustrate that the observed (P,D,M,δ)(P,D,M,\delta)7 pairs lie inside the theoretical region and that higher diversity plus higher shared identity does not reliably predict better performance outcomes (Hoogstra et al., 11 Aug 2025).

The uncertainty problem remains central throughout these analyses. Bayesian fairness models, bootstrap confidence intervals, and posterior-predictive estimates are recurring solutions to the small-(P,D,M,δ)(P,D,M,\delta)8 regime induced by many-way intersections. This suggests a broader QUINTA principle: intersectional measurement is not just about defining the right statistic, but also about quantifying the variance of that statistic under sparsity (Foulds et al., 2018).

6. Governance, reflexivity, and recurring controversies

QUINTA is also a governance vocabulary. In the data-valuation literature, firms can be required to publish an “Information-Leakage Report” for each feature, showing (P,D,M,δ)(P,D,M,\delta)9 and the resulting surcharge; regulators may default to XX0 in the absence of a verifiable estimate; and third-party auditors may be empowered to recompute XX1 using raw logs or synthetic probes. The proposed deployment sites include API billing, privacy dashboards, and session-level consent flows (Garrido-Merchán, 19 Jul 2025).

In AI risk assessment, QUINTA has been operationalized through an LLM-assisted rubric over XX2 reports from XX3 incidents in the AI Incident Database, identifying XX4 harmed subjects with XX5 subject-identification accuracy. At the level of single categories, age and political identity appear at rates comparable to race and gender; at the level of intersections, adolescent girls, lower-class people of color, and upper-class political elites show amplification ratios of XX6, XX7, and XX8, respectively. The framework proposes incorporating such amplification ratios as multipliers in baseline AI risk scores (Bogucka et al., 27 Apr 2026).

A different governance extension makes QUINTA explicitly reflexive. In that formulation, QUINTA is the tuple XX9 over researcher positionality, raw data, methodological parameters, and reflexive metadata, and each pipeline phase—design, collection, cleaning, exploration, modeling, interpretation—requires an explicit mapping that records who is included or excluded, how algorithms may embed or amplify inequality, and what positional stakes shape those choices. The #MeToo demonstration uses this apparatus to contrast a non-QUINTA pathway centered on dominant hashtags with a QUINTA pathway that removes top-degree nodes, raises community entropy from I=(I1,,Im)I=(I_1,\dots,I_m)00 to I=(I1,,Im)I=(I_1,\dots,I_m)01, and surfaces clusters such as #metooblackchurch, #metooqueer, #metoodisabled, and #ustoo (Boyd, 19 Sep 2025).

Several recurring controversies follow from these governance-oriented formulations. One is scope: intersectional harms do not occur only along race and gender, and empirical incident analysis highlights age, class, nationality, and political identity as comparably salient categories in documented AI harms (Bogucka et al., 27 Apr 2026). Another is epistemic status: some frameworks are explicitly model-agnostic, but this does not make them normatively neutral, because parameters such as I=(I1,,Im)I=(I_1,\dots,I_m)02, I=(I1,,Im)I=(I_1,\dots,I_m)03, I=(I1,,Im)I=(I_1,\dots,I_m)04, and I=(I1,,Im)I=(I_1,\dots,I_m)05 encode contested social values directly into pricing or regulation (Garrido-Merchán, 13 Jan 2026). A third is evidentiary: subgroup sparsity, kernel sensitivity, static augmentation under drift, reporting bias in incident datasets, geographical skew, and the absence of causal identification all limit what can be inferred from intersectional statistics alone (Maheshwari et al., 2024, Wang et al., 2022).

In that sense, QUINTA is best understood as a quantitative infrastructure for intersectionality rather than a single doctrine. Its technical repertoire includes mutual information, entropy, PID, hierarchical Bayes, survey-weighted regression, MMD-based synthetic augmentation, clustering, rank-correlation, and risk multipliers; its substantive applications range from digital markets and clinical ML to historical newspapers, mobility surveys, long-COVID prevalence, education, and AI incident analysis; and its governance ambition is to internalize harms, expose compounded disparities, and make explicit the modeling and policy choices through which intersectional categories are quantified (Garrido-Merchán, 19 Jul 2025, Souligne et al., 6 Apr 2026, Varley et al., 2021, Boyd, 19 Sep 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (17)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Quantitative Intersectional Data (QUINTA).