Papers
Topics
Authors
Recent
Search
2000 character limit reached

Alternative routes to universal diversity scaling in component systems: from proteomes to large language models

Published 2 Jul 2026 in cond-mat.stat-mech | (2607.02221v1)

Abstract: Remarkably common statistical laws characterize the diversity scaling and its fluctuations across a wide range of complex "component systems". These regularities are often interpreted as signatures of an underlying innovation mechanism driving the growth of component diversity, but the basic ingredients necessary for their emergence remain poorly understood. In particular, from language and technological artifacts to genomes and gene expression patterns, the number of distinct components grows sublinearly with system size, while its variance scales approximately as the square of its mean. This behavior is consistent across diverse systems, raising the question of whether general constraints or emergent principles underlying diversity and innovation define the architectures of realizations with different numbers of components. To address this question, we derive analytical conditions for the joint emergence of these two diversity laws within a broad class of growth models, showing that they require a specific asymptotic dependence of the innovation probability on diversity and system size. We then demonstrate that the same macroscopic laws arise in a different class of models with latent heterogeneity, where quadratic fluctuation scaling always emerges asymptotically as a consequence of general statistical principles, essentially the law of total variance, without explicitly assuming an innovation mechanism or any specific rule for system assembly. We compare these predictions with empirical data from language, genomes, LEGO constructions, and texts generated by LLMs. Our results show that empirical diversity scaling laws strongly constrain generative models but do not uniquely identify the mechanisms generating diversity, revealing a close correspondence between innovation-driven growth models and latent-variable descriptions.

Summary

  • The paper identifies that a linear diversity-dependent innovation probability (p ∝ h/m) is necessary for reproducing both sublinear Heaps growth and quadratic diversity fluctuations in component systems.
  • It systematically contrasts growth models and latent-variable sampling models, demonstrating that both approaches yield similar macroscopic scaling laws despite distinct underlying mechanisms.
  • Empirical validations in texts, genomes, and LEGO constructions confirm that non-self-averaging behavior is a universal trait of complex component systems.

Universal Diversity Scaling in Component Systems: Routes and Mechanisms

This paper provides a comprehensive analysis of the emergence and universality of diversity scaling laws, specifically Heaps' law and its quadratic fluctuation scaling, in complex component systems. By systematically comparing innovation-driven growth models with latent-variable statistical models, the authors dissect the underlying mechanisms that can lead to the same macroscopic statistical laws across domains ranging from natural language and technological artifacts to genomes and synthetic LLM-generated texts.

Empirical Observations of Diversity Scaling

A broad empirical foundation highlights two key statistical regularities:

  • Heaps’ Law: The number of distinct components (vocabulary size hh) grows sublinearly with system size mm across diverse domains (e.g., texts, LEGO sets, genomes), typically as hmνh \propto m^\nu with 0<ν<10 < \nu < 1.
  • Quadratic Fluctuation Scaling: The variance of diversity across realizations, σh2\sigma_h^2, scales quadratically with its mean, i.e., σh2h2\sigma_h^2 \propto \langle h\rangle^2, revealing strong non-self-averaging behavior and significant cross-realization heterogeneity.

This universality and robustness are demonstrated empirically in Figure 1.

Figure 1

Figure 1: Heaps’ law and quadratic fluctuation scaling in texts, LEGO, Wikipedia, and bacterial genomes, highlighting sublinear vocabulary scaling and variance scaling σh2h2\sigma_h^2 \propto \langle h\rangle^2.

Analytical Growth Models: The Innovation Mechanism

Growth models in the Simon-Yule-CRP-Polya-urn tradition are formulated generically as innovation-duplication stochastic processes parameterized by an innovation probability pn(h,m,j)=αjhβ/mγp_n(h, m, j) = \alpha_j h^\beta / m^\gamma. The critical result is the identification of necessary and sufficient conditions for reproducing both empirical scaling laws:

  • Linear Diversity-Generates-Diversity Mechanism: Only models with an asymptotically linear diversity-dependent innovation probability—specifically pnh/mp_n \sim h/m—yield both sublinear Heaps growth and quadratic diversity fluctuations. This mechanism is necessary; other forms yield different fluctuation scaling Figure 2.

Figure 2

Figure 2: Parameter dependence of diversity scaling and fluctuations in the generic growth model, demonstrating quadratic fluctuation scaling only for parameters yielding pnh/mp_n \sim h/m.

Analytical calculations confirm that for mm0, both mm1 and mm2 emerge in the asymptotic regime.

Validation in LLMs

Empirical analyses in both human and LLM-generated texts show that the effective innovation rate is indeed linearly proportional to current diversity, confirming that autoregressive LLMs effectively implement the diversity-generates-diversity mechanism Figure 3.

Figure 3

Figure 3: Empirical innovation probability vs. current diversity in human and LLM-generated text, demonstrating linear dependence as predicted by the growth model.

Latent-Variable Sampling Models: Heterogeneity Without Explicit Innovation

As an alternative, the paper investigates sampling models with latent variables (e.g., topics, functions, taxonomic clades). In this latent heterogeneity framework, each realization is assembled according to a (possibly topic-dependent) component frequency distribution.

  • Law of Total Variance: Quadratic fluctuation scaling emerges generally as a statistical consequence of averaging over realizations with different underlying component frequencies, without requiring diversity-generating innovation.
  • Generative Equivalence: The latent-variable model produces the same phenomenology (sublinear Heaps law, quadratic scaling for mm3, and inferred mm4) as the innovation-driven growth model for sufficiently strong heterogeneity, as illustrated by simulation of a Dirichlet-topic sampling model Figure 4.

Figure 4

Figure 4: Sampling model with Dirichlet topic heterogeneity transitions from Poisson to quadratic fluctuation scaling as topic variance increases.

Empirical Case Studies: Genomes and LEGO Constructions

The analysis of annotated datasets such as bacterial genomes (taxonomic clades as topics) and LEGO constructions (themes as topics) confirms that the diversity scaling within annotated groups obeys power-law Heaps' laws with different prefactors. However, conditioning on known topics neither erases quadratic fluctuation scaling nor fully accounts for all observed diversity, indicating the existence of additional, possibly hierarchical, latent heterogeneity Figure 5.

Figure 5

Figure 5: Topic-conditioned Heaps’ laws and variance scaling in genomes and LEGO constructions; topic labels induce vertical shifts in vocabulary scaling, but quadratic fluctuations persist globally.

Theoretical Implications and Phenomenological Equivalence

  • Non-uniqueness of Mechanistic Inference: The same macroscopic statistical laws may originate from fundamentally distinct micro-mechanisms—growth by innovation with feedback or sampling with latent heterogeneity.
  • Exchangeability and Statistical Representation: The equivalence of these generative models at the level of vocabulary statistics is reminiscent of results like de Finetti’s theorem for exchangeable processes.
  • Non-self-averaging: Both frameworks predict non-self-averaging behavior for diversity: relative fluctuations remain finite even as system size diverges, contrary to classical sampling models.

Implications and Future Directions

The findings have significant implications for the interpretation of macroscopic diversity scaling across fields such as linguistics, molecular biology, technology, and artificial intelligence:

  • Caution in Mechanistic Attribution: Empirical observation of Heaps’ law and quadratic fluctuation scaling does not uniquely evidence an innovation-driven process; latent variable heterogeneity suffices.
  • Constraint on Generative Modeling: Generative models (especially stochastic generative LLMs) must conform to observed diversity scaling, tightly constraining permissible mechanisms.
  • Future Discriminative Observables: Discrimination between mechanisms may require observables beyond marginal statistics—e.g., higher-order correlations, temporal dynamics, and full characterization of latent structure.

Conclusion

This work rigorously demonstrates that universal diversity scaling laws in component systems—namely, sublinear growth and non-self-averaging quadratic fluctuations—arise from either explicit diversity-dependent innovation mechanisms or latent-variable heterogeneity. The phenomenological equivalence of these mechanisms calls for caution in causal interpretation of scaling laws and points toward a statistical universality class encompassing both. Identifying discriminative empirical or theoretical signatures capable of distinguishing these routes remains an important open problem for complexity science, statistical mechanics, and generative modeling.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.