Papers
Topics
Authors
Recent
Search
2000 character limit reached

Testing for Single-Population Ancestry in the Admixture Model

Published 1 Jun 2026 in stat.ME and math.ST | (2606.01990v1)

Abstract: The Admixture Model describes genetic marker data by representing each individual's genome as a mixture of contributions from KK ancestral populations, with the individual admixture vector summarizing the corresponding ancestry proportions. In population and forensic genetics, a key question is whether an individual's genome supports a predominantly single-ancestry interpretation or whether an admixed interpretation is more appropriate. We propose a statistical test for single-population ancestry in the supervised Admixture Model, where ancestral allele frequencies are treated as known. The test assesses whether the largest admixture component exceeds a practitioner-chosen dominance threshold, giving precise meaning to the notion of a sufficiently strong single-population contribution. To calibrate the test, we develop a constrained parametric bootstrap procedure that generates data under a null-constrained maximum likelihood estimator, accounting for the constrained hypothesis structure, the marker-wise heterogeneity and small sample sizes. Under standard regularity conditions, we prove that the proposed test has asymptotic level αα and is consistent, ensuring control of false single-ancestry declarations while reliably detecting dominant ancestry components. Simulation studies demonstrate good finite-sample performance across different numbers of ancestral populations, marker-panel sizes, dominance thresholds, and allele-frequency distributions. We further illustrate the practical utility of the method using data from the 1000 Genomes Project. The proposed framework delivers interpretable, threshold-based ancestry assessment with rigorous error control, and extends constrained bootstrap methodology to the independent but non-identically distributed setting of genetic marker data.

Summary

  • The paper develops a constrained parametric bootstrap test that evaluates whether an individual’s largest ancestry component exceeds a chosen dominance threshold, with asymptotic level control and consistency.
  • Simulations show that type I error approaches the nominal 5% level at the boundary and that power rises with marker count, reaching full power in some settings by 1,000 markers.
  • Applied to 55 markers from the 1000 Genomes Project, the test rejects single-population ancestry for 91% of AFR, 83% of EAS, 57% of EUR, 41% of SAS, and 13% of AMR individuals, while results remain limited by estimated reference frequencies and admixed populations.

This paper develops a formal hypothesis test for single-population ancestry within the supervised Admixture Model, addressing an inferential gap left by prior work that focused exclusively on estimation of individual admixture proportions. The authors propose a constrained parametric bootstrap test of whether the largest admixture component exceeds a practitioner-specified dominance threshold, prove asymptotic level control and consistency, and validate the procedure through simulation and an application to 1000 Genomes Project data (2606.01990).

Model and problem setting

The setup assumes MM markers in linkage equilibrium, K2K \geq 2 non-admixed ancestral populations with known allele frequencies pkimp_{kim}, and a diploid donor whose genotype at marker mm is multinomial with $2$ trials and probabilities given by the mixture q,pim\langle q', p_{\cdot im} \rangle, where qSKq' \in \mathbb{S}^K is the individual admixture vector. The parameter space is restricted to a compact interior set Θ\Theta bounded away from the simplex boundary by a constant εq\varepsilon_q (assumption (A1)), a standard device to ensure regular maximum likelihood asymptotics and to rule out non-standard limiting behavior of the MLE.

The estimand is the dominant component d=maxkqkd_\infty = \max_k q'_k, and the test contrasts K2K \geq 20 against K2K \geq 21 for a threshold K2K \geq 22. The threshold interpretation is deliberately genealogical rather than pedigree-forensic: K2K \geq 23 corresponds to the expected allele contribution of three grandparents from the same ancestral population, and K2K \geq 24 to seven great-grandparents. The paper is explicit that the test does not infer actual pedigree, and that exact purity (K2K \geq 25) lies on the boundary outside the scope of the asymptotics but can be approximated within K2K \geq 26.

The constrained parametric bootstrap test

The test statistic is K2K \geq 27 from the MLE K2K \geq 28. Calibration proceeds via a null-constrained estimator: if K2K \geq 29 the unconstrained MLE is used; otherwise the constrained MLE pkimp_{kim}0, maximizing the likelihood subject to pkimp_{kim}1, is used. Bootstrap data are generated from this null-constrained parameter, the bootstrap MLE and statistic are computed over pkimp_{kim}2 replicates, and pkimp_{kim}3 is rejected when pkimp_{kim}4 exceeds the empirical pkimp_{kim}5-quantile of the bootstrap distribution. The paper argues that a Wald-type alternative would require estimating the variance of a maximum of correlated components and would rely on a normal approximation that may be poor for the small marker panels (pkimp_{kim}6) typical of forensic practice; the bootstrap calibrates the least-favorable boundary case directly.

Two extensions are provided. First, a swapped hypothesis pair pkimp_{kim}7 versus pkimp_{kim}8 controls the error of falsely concluding admixed ancestry. Second, a sequential testing procedure over a decreasing grid of thresholds, using the sequential rejection principle, returns the largest grid threshold for which the data support single-population dominance, with asymptotic level pkimp_{kim}9 preserved by monotonicity of the adjusted critical values.

Asymptotic theory

The theoretical results rest on four assumptions beyond (A1): a uniform lower bound mm0 on all allele frequencies (A2); weak convergence of the empirical marker design measure mm1 to an asymptotic design mm2 (A3); and an identifiability condition (A4) requiring that the design be sensitive to all directions mm3 orthogonal to the all-ones vector. The paper notes that (A4) implies positive Kullback–Leibler divergence between any two distinct admixtures, i.e., the marker panel must contain enough ancestry-informative variation; if ancestral populations had identical allele frequencies across markers, the test could not distinguish admixtures.

The main theorem establishes, conditionally on the data, that mm4 converges to mm5 with mm6 the limiting Fisher information, and that the test has rejection probability tending to mm7 for mm8, exactly mm9 on the boundary $2$0, and $2$1 under the alternative. The proof strategy combines Hoadley's asymptotics for MLEs under independent but non-identically distributed observations with the constrained bootstrap framework of Dette and Möllenhoff. A methodological contribution claimed by the authors is that the constrained bootstrap theory is here extended to the independent, non-identically distributed setting induced by marker-wise heterogeneity in allele frequencies.

Finite-sample performance

Simulations vary $2$2, $2$3, thresholds near $2$4–$2$5 and $2$6, and allele-frequency distributions (Dirichlet(1,1) versus Dirichlet(0.5,2)), with $2$7 and $2$8; increasing $2$9 to 1000 was found not to materially change results. Inside the null, type I error is close to zero across all q,pim\langle q', p_{\cdot im} \rangle0, and at the boundary it converges to the nominal level, with mild liberalism at small q,pim\langle q', p_{\cdot im} \rangle1 (rejection rate q,pim\langle q', p_{\cdot im} \rangle2 at q,pim\langle q', p_{\cdot im} \rangle3, q,pim\langle q', p_{\cdot im} \rangle4, q,pim\langle q', p_{\cdot im} \rangle5) and conservatism in the hardest setting (q,pim\langle q', p_{\cdot im} \rangle6, q,pim\langle q', p_{\cdot im} \rangle7, rejection rate q,pim\langle q', p_{\cdot im} \rangle8 at q,pim\langle q', p_{\cdot im} \rangle9). Power increases with qSKq' \in \mathbb{S}^K0 and with extremity of the alternative; for qSKq' \in \mathbb{S}^K1, qSKq' \in \mathbb{S}^K2, full power is reached at qSKq' \in \mathbb{S}^K3 already at qSKq' \in \mathbb{S}^K4.

The comparison across qSKq' \in \mathbb{S}^K5 is instructive and somewhat counterintuitive: the test performs better for qSKq' \in \mathbb{S}^K6 than qSKq' \in \mathbb{S}^K7 at threshold qSKq' \in \mathbb{S}^K8, but better for qSKq' \in \mathbb{S}^K9 at threshold Θ\Theta0. The authors attribute this to the geometry of the simplex—at the higher threshold with Θ\Theta1 the residual mass Θ\Theta2 is spread over several small components, placing Θ\Theta3 near the boundary, making constrained estimation difficult and producing conservative critical values and reduced power. This indicates that finite-sample performance depends on the position of the true admixture within the simplex, not only on Θ\Theta4, Θ\Theta5, and the threshold. More concentrated allele frequencies (the Beta(0.5,2) setting) reduce power relative to uniform frequencies, as expected from reduced marker informativeness.

Application to the 1000 Genomes Project

Using the 55 bi-allelic marker panel of Kidd et al. and the five 1000 Genomes superpopulations as references (Θ\Theta6, threshold Θ\Theta7, Θ\Theta8), the test rejects single-population ancestry for approximately Θ\Theta9 of AFR and εq\varepsilon_q0 of EAS individuals, roughly half of EUR (εq\varepsilon_q1) and SAS (εq\varepsilon_q2) individuals, and only εq\varepsilon_q3 of AMR individuals. Under the swapped hypothesis with threshold εq\varepsilon_q4, εq\varepsilon_q5 of the full data set is rejected, with strong heterogeneity: about εq\varepsilon_q6 for AMR versus εq\varepsilon_q7 for EAS. These patterns are consistent with AMR being itself an admixed and heterogeneous reference group, while AFR and EAS are more strongly differentiated by the marker set.

The application departs from the theoretical framework in two ways that the authors acknowledge: the reference allele frequencies are estimated from the data (leave-one-individual-out) rather than known, and the superpopulations—particularly AMR—are not truly non-admixed ancestral populations. The exercise is therefore an illustration rather than a validation of the asymptotic guarantees in a realistic setting.

Limitations and open questions

Several limitations are stated plainly. The theory assumes known ancestral allele frequencies; extending validity to estimated reference frequencies is left as an open problem, and this is precisely the setting of the data application. The parameter space excludes the simplex boundary, so exact pure ancestry cannot be handled by the asymptotic theory. The dominance threshold is application-dependent and must be justified externally; the sequential grid procedure mitigates but does not eliminate this dependence. The linkage-equilibrium assumption excludes linked-marker models, and extension to such models would require substantially different techniques. Finally, the paper leaves open the construction of analogous tests for the minimum ancestry component, which would be relevant to the unsupervised model and could provide a principled inferential approach to selecting the number of ancestral populations εq\varepsilon_q8 in software such as Structure and ADMIXTURE.

Conclusion

The paper provides a statistically rigorous, threshold-based decision procedure for single-population ancestry that complements existing estimation theory for the Admixture Model. Its main contributions are a constrained parametric bootstrap calibrated at the least-favorable boundary, asymptotic level and consistency guarantees under independent but non-identically distributed markers, and evidence of adequate finite-sample behavior on panels of realistic forensic size. The guarantees, however, are conditional on known reference allele frequencies and interior parameter values, and closing the gap between the theoretical setting and the estimated-frequency setting used in practice remains the central open question raised by this work.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.