Quantitative Intersectional Data (QUINTA)
- QUINTA is a set of quantitative frameworks that jointly model intersecting social identities to measure and govern multiple protected attributes.
- It employs information-theoretic metrics like mutual information to price informational harm and assess data leakage in a model-agnostic manner.
- QUINTA supports fairness auditing in machine learning by constructing joint distributions over protected attributes to identify and mitigate intersectional biases.
Searching arXiv for the cited QUINTA-related papers and closely related intersectional methodology papers. arxiv_search: (Garrido-Merchán, 19 Jul 2025) arxiv_search: (Garrido-Merchán, 13 Jan 2026) arxiv_search: (Souligne et al., 6 Apr 2026) arxiv_search: (Morina et al., 2019) arxiv_search: (Foulds et al., 2018) arxiv_search: (Varley et al., 2021) Quantitative Intersectional Data (QUINTA) denotes a family of quantitative frameworks for representing, measuring, and governing data about intersecting social identities. In recent work, the acronym is expanded in several ways—“QUantitative INTersectional data lAbeling and valuation,” “QUANTitative INtersectionAl data,” “QUantitative INtersectional TAxonomy,” and “QUANTitative InTersectional Analysis”—but the common premise is consistent: protected attributes such as race, gender, disability, class, sexuality, age, education, or political identity are modeled jointly, and the relevant quantities are defined on the joint distribution or on explicitly constructed intersectional groups rather than on single axes in isolation (Garrido-Merchán, 19 Jul 2025, Garrido-Merchán, 13 Jan 2026, Souligne et al., 6 Apr 2026, Bogucka et al., 27 Apr 2026).
1. Terminological range and analytical object
Across the literature, QUINTA is not a single algorithm. It is a recurrent label for quantitative intersectional workflows in data valuation, fairness auditing, survey analysis, historical NLP, educational measurement, incident analysis, and reflexive AI research. The shared analytical move is to define an intersectional object explicitly: a joint protected profile such as or , a Cartesian product of group axes , or a finite set of aggregate identities with proportions (Garrido-Merchán, 19 Jul 2025, Garrido-Merchán, 13 Jan 2026, Hoogstra et al., 11 Aug 2025).
| Use of QUINTA | Primary object | Representative paper |
|---|---|---|
| Data valuation | Joint protected profile or | (Garrido-Merchán, 19 Jul 2025, Garrido-Merchán, 13 Jan 2026) |
| Fairness auditing | Intersectional groups from multiple protected attributes | (Souligne et al., 6 Apr 2026, Morina et al., 2019) |
| Underrepresentation correction | Sensitive groups and their intersections | (Diana et al., 2023) |
| Empirical social analysis | Joint strata over gender, age, education, race, sexuality, or class | (Vallée et al., 2021, Cohen et al., 3 Mar 2026) |
| Reflexive AI/DS practice | Tuple over positionality, data, methods, and reflexive metadata | (Boyd, 19 Sep 2025) |
This terminological variation matters because the literature uses “QUINTA” both narrowly and broadly. In some papers it names a formal pricing rule for the social cost of intersectional inference; in others it names an auditing toolkit, a survey-analysis protocol, or a reflexive methodological paradigm. Taken together, these uses suggest that QUINTA is best understood as a quantitative intersectional methodology class rather than as a single standardized framework (Boyd, 19 Sep 2025).
2. Information-theoretic valuation and Pigouvian pricing
A central QUINTA line of work formulates intersectional privacy as an externality. The core claim is that the “social harm” of releasing a datum derives exclusively from how much it reduces uncertainty about a full protected profile 0 taken jointly. The corresponding metric is Shannon mutual information,
1
which quantifies the expected drop in entropy about the protected profile when 2 is observed (Garrido-Merchán, 19 Jul 2025).
Under the axioms Intersectional Externality, No-Harm/No-Cost, Monotonicity, and Additivity, the price of a datum is forced into the linear form
3
with pure surcharge
4
where 5 is the data-collection cost and 6 is the regulator’s social-cost parameter per unit of information. A related pointwise formulation prices each observed data point by
7
thereby translating informational harm into dollars per nat or per bit (Garrido-Merchán, 19 Jul 2025, Garrido-Merchán, 13 Jan 2026).
These formulations are explicitly model-agnostic. The surcharge depends on 8 or on the posterior/prior pair 9 and 0, not on any specific downstream predictive model. The same rule can therefore be implemented whether the joint distribution is estimated via histograms, kernel-density estimation, k-nearest-neighbour estimators, logistic regression, random forests, auto-encoders, GANs, mixture models, or variational bounds such as MINE (Garrido-Merchán, 19 Jul 2025, Garrido-Merchán, 13 Jan 2026).
The operational recipe is straightforward. One discretizes or otherwise models the joint intersectional space, collects an audited sample of 1 or 2, estimates the joint and conditional distributions, computes 3, and multiplies by a policy parameter. The papers propose calibration by setting a statutory maximum penalty 4 for full disclosure and choosing 5. They also give concrete examples: for a race 6 gender 7 disability space with 8 joint categories, 9 nats under a uniform prior; a feature leaking 0 nats would incur a surcharge of \$c\in\mathcal C$1\lambda=100{,}000\ \$c\in\mathcal C$2, whereas a timestamp leaking $c\in\mathcal C$3 nats would cost about \$2,000 (Garrido-Merchán, 19 Jul 2025).
A further extension introduces normative weights 4 on intersectional groups, yielding a weighted mutual information
5
with 6. This makes the pricing rule not merely model-agnostic but explicitly norm-sensitive: 7 and 8 become the policy knobs through which contested social values and redistributive aims are embedded into market pricing (Garrido-Merchán, 13 Jan 2026).
3. Fairness auditing, estimation, and mitigation in machine learning
A second major QUINTA strand addresses intersectional fairness in classification. The basic setup constructs the full intersectional domain as the Cartesian product of protected attributes and measures disparities on joint subgroups rather than on marginal attributes alone. Formal criteria include intersectional statistical parity, equal opportunity, false-positive parity, equalized odds, and multiplicative notions such as 9-differential fairness, where the worst-case log-ratio between subgroup probabilities is bounded by 0 (Morina et al., 2019).
Several papers emphasize that intersectional estimation is statistically difficult because data sparsity grows rapidly with the number of protected attributes and categories. Proposed responses include beta-binomial smoothing, bootstrap confidence intervals, Bayesian posterior estimation, and hierarchical Bayesian models that shrink sparse group estimates toward a global predictor. In the Bayesian setting, posterior samples of group-conditional probabilities induce posterior samples of fairness metrics such as differential fairness and subgroup fairness, producing uncertainty-aware estimates rather than unstable plug-in ratios (Foulds et al., 2018, Morina et al., 2019).
The clinical toolkit FairLogue operationalizes this logic in both observational and counterfactual settings. Its observational framework extends demographic parity, equalized odds, and equal opportunity difference to intersectional groups, while its generalized counterfactual framework evaluates fairness under interventions on group membership and compares observed disparities to permutation-based null distributions through a “u-value.” In a glaucoma surgery prediction task on the All of Us Controlled Tier V8 cohort, observational analysis found an intersectional demographic parity gap of 1, an equalized-odds true positive rate gap of 2, and a false positive rate gap of 3, all larger than the corresponding single-axis gaps; the counterfactual analysis, however, produced u-values near 4, suggesting that the observed disparities were consistent with chance after conditioning on covariates (Souligne et al., 6 Apr 2026).
Mitigation methods span post-processing, reweighting, and data augmentation. One post-processing framework constructs a randomized thresholding derived predictor with subgroup-specific thresholds and flip probabilities, then solves a linear program to satisfy chosen fairness constraints with minimal expected loss (Morina et al., 2019). For underrepresentation bias, another QUINTA formulation assumes a small unbiased sample and a larger biased sample in which positive examples are filtered at unknown group-wise drop-out rates 5. It estimates the 6, defines a reweighting function on positive examples, and proves agnostic PAC learnability under intersectional bias for finite-VC hypothesis classes (Diana et al., 2023).
Synthetic augmentation is handled differently in “Synthetic Data Generation for Intersectional Fairness by Leveraging Hierarchical Group Structure,” which treats intersectional groups as intersections of parent categories. The method learns a generative transformation from immediate parent groups to scarce child groups using an MMD objective and then trains downstream classifiers under equal sampling across groups. On four datasets, the reported empirical pattern is a Balanced Accuracy drop of at most 7 points versus the unconstrained baseline, worst-case FPR reductions of up to 8, reductions in differential fairness on Hate Speech from about 9 to 0, reductions in 1 from about 2 to 3 on Anxiety and from about 4 to 5 on Twitter Hate Speech, and no method-induced “leveling down” (Maheshwari et al., 2024).
A recurrent caution is that handling underrepresentation is not purely technical. Reweighting, resampling, GAN-based augmentation, and counterfactual augmentation all have normative implications, and one paper explicitly warns that synthetic face-image generation can reproduce disturbing tropes analogous to historical “face painting,” while unsupervised relabeling of sensitive categories can undermine self-identification and agency (Wang et al., 2022).
4. QUINTA as an empirical workflow across domains
Outside ML fairness, QUINTA appears as a general empirical protocol for intersectional measurement. In historical NLP, it combines OCR correction, tokenization, distributional semantics, and bias metrics. “Measuring Intersectional Biases in Historical Documents” uses Pointwise Mutual Information, WEAT effect sizes, and a two-dimensional mapping of descriptors into a gender 6 race plane. The study reports that BPE yields average nearest-neighbour Jaccard similarity about 7 but only about 8 misspelling compatibility, whereas spaCy preprocessing yields about 9 stability and about 0 compatibility; it also reports that white females are strongly associated with “family,” “weak,” and “appearance,” white males with “career,” “strong,” and “manual labour,” and that more than 1 of strongly “beautiful” adjectives fall into the white-female quadrant only (Borenstein et al., 2023).
In urban geography, QUINTA is used to study hourly spatial distributions across gender, age, and education. The French mobility-survey application relies on 2 origin–destination surveys, 3 respondents, 4 trips, 5 districts, and 6 hourly snapshots. It normalizes district-level group counts into probability-like temporal profiles, computes a Euclidean-based mismatch index, and applies Ward clustering. The reported results include five district temporal profiles, gender-based mismatch in 7 of districts, at least one age-based mismatch in 8, at least one education-based mismatch in 9, and a synchronization contrast in which dominant groups are more synchronous than non-dominant groups in high-flux districts but not in stable districts (Vallée et al., 2021).
In population health, the long-COVID study operationalizes intersectional strata over sex, race/ethnicity, education, and sexual and gender minority status, then combines weighted descriptive prevalence estimates with survey-corrected logistic regressions containing two-way and three-way interactions. The descriptive results reveal a “step-like pattern” in which college-educated men have the lowest prevalence of long COVID and women without college educations the highest, while the regression results show, among other effects, odds ratios of 0 for Female, 1 for BA or higher, 2 for SGM, and an interaction odds ratio of 3 for Female 4 Hispanic in the long-COVID model (Cohen et al., 3 Mar 2026).
Educational research uses a related but conceptually distinct QUINTA under the label of critical quantitative intersectionality. The physics-learning study fits hierarchical linear models with race, gender, and course-type effects, and then evaluates equity through two operationalizations: Equality of Learning and Equity of Individuality. The reported result is deliberately non-convergent: collaborative instruction improved gains for all groups by about 5 percentage points, satisfying Equity of Individuality, but the gaps relative to White men remained, so Equality of Learning was not achieved (Dusen et al., 2018).
A more recent evaluation framework for LLM empathy uses controlled prompt generation over intersectional social categories such as race, pronoun, sexuality, education, religion, age, and socioeconomic status, then measures Cognitive Empathy by five-way multiple-choice accuracy, affective empathy through VAD scores, and response appropriateness through Epitome-derived scores. In the initial evaluation sample, intersectional cell variances were all below 6, no cell mean differed significantly from the grand mean after correction, and one outlier cell—Lesbian plus “she”—showed an interpretation score about 7 standard deviations above the overall mean (Formanek et al., 2024).
5. Higher-order structure: synergy, diversity, and uncertainty
Some QUINTA frameworks move beyond groupwise disparity measurement to quantify higher-order structure directly. Partial Information Decomposition is used to separate redundant, unique, and synergistic contributions of intersecting identities to an outcome. For two predictors 8 and target 9,
0
This formalism is used to argue that some intersectional effects are irreducible to single identities considered individually. In U.S. census-style analyses, the reported decomposition of 1 yields about 2 redundant information, 3 synergistic information, and 4 unique information, while a synthetic-data comparison shows that standard linear regression with multiplicative interaction terms fails to distinguish truly synergistic effects from redundant ones (Varley et al., 2021).
Another line defines multidimensional diversity metrics directly on aggregate identities. Intersecting Diversity,
5
is a normalized probability that two random draws differ in at least one trait. Shared Identity,
6
is the expected fraction of traits shared by two draws. These measures are shown to be anti-correlated, with admissible pairs constrained to a polygonal region and with no single point that maximizes both metrics simultaneously. The case studies on Hollywood films, Survivor tribes, and North American companies illustrate that the observed 7 pairs lie inside the theoretical region and that higher diversity plus higher shared identity does not reliably predict better performance outcomes (Hoogstra et al., 11 Aug 2025).
The uncertainty problem remains central throughout these analyses. Bayesian fairness models, bootstrap confidence intervals, and posterior-predictive estimates are recurring solutions to the small-8 regime induced by many-way intersections. This suggests a broader QUINTA principle: intersectional measurement is not just about defining the right statistic, but also about quantifying the variance of that statistic under sparsity (Foulds et al., 2018).
6. Governance, reflexivity, and recurring controversies
QUINTA is also a governance vocabulary. In the data-valuation literature, firms can be required to publish an “Information-Leakage Report” for each feature, showing 9 and the resulting surcharge; regulators may default to 0 in the absence of a verifiable estimate; and third-party auditors may be empowered to recompute 1 using raw logs or synthetic probes. The proposed deployment sites include API billing, privacy dashboards, and session-level consent flows (Garrido-Merchán, 19 Jul 2025).
In AI risk assessment, QUINTA has been operationalized through an LLM-assisted rubric over 2 reports from 3 incidents in the AI Incident Database, identifying 4 harmed subjects with 5 subject-identification accuracy. At the level of single categories, age and political identity appear at rates comparable to race and gender; at the level of intersections, adolescent girls, lower-class people of color, and upper-class political elites show amplification ratios of 6, 7, and 8, respectively. The framework proposes incorporating such amplification ratios as multipliers in baseline AI risk scores (Bogucka et al., 27 Apr 2026).
A different governance extension makes QUINTA explicitly reflexive. In that formulation, QUINTA is the tuple 9 over researcher positionality, raw data, methodological parameters, and reflexive metadata, and each pipeline phase—design, collection, cleaning, exploration, modeling, interpretation—requires an explicit mapping that records who is included or excluded, how algorithms may embed or amplify inequality, and what positional stakes shape those choices. The #MeToo demonstration uses this apparatus to contrast a non-QUINTA pathway centered on dominant hashtags with a QUINTA pathway that removes top-degree nodes, raises community entropy from 00 to 01, and surfaces clusters such as #metooblackchurch, #metooqueer, #metoodisabled, and #ustoo (Boyd, 19 Sep 2025).
Several recurring controversies follow from these governance-oriented formulations. One is scope: intersectional harms do not occur only along race and gender, and empirical incident analysis highlights age, class, nationality, and political identity as comparably salient categories in documented AI harms (Bogucka et al., 27 Apr 2026). Another is epistemic status: some frameworks are explicitly model-agnostic, but this does not make them normatively neutral, because parameters such as 02, 03, 04, and 05 encode contested social values directly into pricing or regulation (Garrido-Merchán, 13 Jan 2026). A third is evidentiary: subgroup sparsity, kernel sensitivity, static augmentation under drift, reporting bias in incident datasets, geographical skew, and the absence of causal identification all limit what can be inferred from intersectional statistics alone (Maheshwari et al., 2024, Wang et al., 2022).
In that sense, QUINTA is best understood as a quantitative infrastructure for intersectionality rather than a single doctrine. Its technical repertoire includes mutual information, entropy, PID, hierarchical Bayes, survey-weighted regression, MMD-based synthetic augmentation, clustering, rank-correlation, and risk multipliers; its substantive applications range from digital markets and clinical ML to historical newspapers, mobility surveys, long-COVID prevalence, education, and AI incident analysis; and its governance ambition is to internalize harms, expose compounded disparities, and make explicit the modeling and policy choices through which intersectional categories are quantified (Garrido-Merchán, 19 Jul 2025, Souligne et al., 6 Apr 2026, Varley et al., 2021, Boyd, 19 Sep 2025).