- The paper introduces admixed arrays, paired binary matrices encoding ancestry and allele data, and rigorously derives their asymptotic enumeration under biological constraints.
- It employs generating functions and saddle-point approximations to obtain precise expansions for doubly constrained arrays, including explicit fourth-moment correction terms.
- The study reveals that classical independence heuristics overcount configuration sizes, necessitating an analytic correction factor due to negative correlation of constraints.
Asymptotic Enumeration of Admixed Arrays and an Alternative Independence Heuristic
This work introduces and rigorously analyzes admixed arrays, a class of paired binary matrices motivated by the structure of large-scale genetic data in population genetics. Admixed arrays encode ancestry and allele information for N diploid individuals across P loci, linking them to biological applications such as local ancestry inference and estimation of ancestry-specific allele frequencies. Formally, each admixed array consists of two N×2P binary matrices [A,X], where A encodes local ancestries and X encodes corresponding allele dosages. The combinatorial complexity arises from the pairing of columns (grouping homologous chromosomes by locus) and the introduction of biological constraints, leading to new combinatorial objects distinct from classical contingency tables or bipartite incidence structures.
Two primary families of constraints on these arrays are considered:
- Row-sum Constraint (“global ancestry”): Each individual's cumulative ancestry (row sum) is prescribed.
- Paired Column-sum Constraint (“ancestry-specific dosage”): At each locus, only subsets of possible paired column sums are feasible, encoding the joint distribution of ancestries and alleles—a nontrivial inequality constraint unique to admixed arrays.
Both constraints reflect quantities of direct interest in statistical genetics, such as controlling for population structure and estimating local ancestry-specific allele frequencies.
Figure 1: Example of a two-way admixed array depicting the pairing of alleles (columns), individual ancestry states (rows/colors), and the inheritance of local ancestry and allele dosage information as used in this work.
Exact Enumeration and Entropic Comparison of Constraint Families
Precise formulas are derived for both single-constraint cases A1 (row-sum only) and A2 (paired column-sum only). For A1, enumeration reduces to a product of binomials over all rows times unconstrained choices for X. For A2, a more delicate combinatorial analysis accounts for the feasible set of paired column sums, reflecting minimal coupling even in unconstrained A. Notably, this introduces combinatorial feasibility inequalities not present in classical binary matrix theory.
The relative restrictiveness of the two constraint families is characterized by an information-theoretic analysis: the log-cardinality of each family admits a sharp entropy approximation in the dense, semi-regular regime (uniform margins and densities bounded away from extremal values). In this regime, the following criterion holds with arbitrarily small error for large [A,X]0: [A,X]1 is larger than [A,X]2 if and only if
[A,X]3
where [A,X]4 is the mean binary entropy for the row-constraint and [A,X]5 is the mean ternary entropy (reflecting the count of the three ancestry/dosage states per column), and [A,X]6 is the mean sum of ancestry-specific allele fractions.
Figure 2: Visual agreement between the entropy-based approximate criterion and ground-truth enumeration for the comparison of constraint restrictiveness. Misclassification is concentrated where entropy differences are near-zero.
The fraction of parameter space where the entropy criterion and exact counting disagree decays as [A,X]7 in the semi-regular case. Notably, the paired column-sum constraints induce a departure from the classical Gale–Ryser setting due to coupled inequalities and combinatorial shadowing effects.
Asymptotic Enumeration: Doubly Constrained Admixed Arrays
The enumeration of doubly constrained admixed arrays ([A,X]8) subject to both global ancestry and ancestry-specific allele dosage constraints presents a fundamentally new technical challenge, with the feasible region displaying significant coupling of constraints over the full bipartite pairing.
Using generating function techniques and saddle-point approximation, the paper derives a precise asymptotic expansion for the log-cardinality in the semi-regular ([A,X]9 margin) case:
A0
with A1. The expansion isolates the explicit fourth-moment correction and demonstrates rigorous control of the remainder via probabilistic and hypercontractivity techniques.
The proof constructs a Laurent generating function encoding all constraints as polynomial weights, proceeds via dimensionality reduction exploiting torus invariance and Cauchy’s integral theorem, and then evaluates the resulting high-dimensional Gaussian integral, controlling higher-order contributions.
Figure 3: Visualization of the entropy function, feasible region sizing, and the geometric structure used for bounding error rates and saddle-point concentration.
Correction to the Independence Heuristic
A highlight of the theory is its contradiction of the classical independence heuristic for jointly constrained binary or integer matrices. For standard dense, semi-regular binary or integer contingency tables, the heuristic provides that
A2
with a dimension-free correction factor A3 depending on parity (with A4 the normalization factor), reflecting weak dependence of row and column events.
In contrast, for admixed arrays in the regime A5 with moderate density, an analytic correction is necessary: the independence heuristic overcounts by a factor of A6,
A7
This signifies that row and paired column constraints in admixed arrays are asymptotically negatively correlated, even under semi-regularity and density constraints, and that the classical correction fails. The quantitative deviation stems from structural coupling induced by biological pairing and joint constraint feasibility, and is directly visible in the analytic expansion.
Numerical Validation and Algorithmic Contributions
Exact enumeration for moderate A8 is made feasible by an algorithmic extension of the Miller & Harrison dynamic programming algorithm, leveraging memoization and parallelization on conjugate vectors. Extensive simulations confirm tight agreement between saddle-point estimates and exact counts, and numerically demonstrate rapid convergence of the independence heuristic correction factor to the predicted value as dimensions grow.
Implications, Extensions, and Future Directions
These results expand the taxonomy of constrained matrix models, demonstrating that even modest structural innovations (here, pairing and paired constraints motivated by genetics) can yield genuinely new corrections to fundamental probabilistic heuristics underlying high-dimensional combinatorics.
Practical implications include more accurate quantification of configuration space sizes for constrained genetic data simulation, calibration of permutation-based tests, and the robust interpretation of genetic summary statistics under complex dependency. Theoretically, the work suggests that for other classes of constraints arising in data-informed combinatorics, subtle but explicit corrections to classical independence heuristics (including the value of the correction factor) may arise in dense regimes.
Potential extensions include:
- Analytic and computational study of higher-way (A9) admixed arrays, where the algebraic and geometric coupling grows,
- Asymptotic enumeration under alternative scaling limits (X0, sparse regime),
- Broader application to colored or weighted bipartite models with paired or “linked” constraints,
- Explicit enumeration for contingency tables arising from other biological or network models with nontrivial marginal or structural dependencies.
Conclusion
This work provides a detailed analytic and algorithmic framework for counting admixed arrays under biologically motivated coupled constraints, establishing new asymptotic phenomena in the structure of high-dimensional discrete models and demonstrating the necessity of model-specific independence heuristics. The deviation from the canonical X1 correction is both quantitatively explicit and biologically interpretable, and the technical methods (entropic comparison, saddle-point analysis, hypercontractivity) are broadly transferrable to other dependent combinatorial enumeration problems.