Papers
Topics
Authors
Recent
Search
2000 character limit reached

PU-Lie Models in Phylogenetics

Updated 3 July 2026
  • PU-Lie models are Lie Markov models defined by purine/pyrimidine symmetry, ensuring closure under matrix exponentiation and log-composition.
  • They extend classical models like Jukes–Cantor, Kimura, and HKY by using Lie algebraic properties to capture time-inhomogeneous evolutionary processes.
  • The framework unifies nucleotide substitution modeling with enhanced parameter identifiability and efficient likelihood optimization for phylogenetic inference.

PU-Lie refers to a family of Lie Markov models with purine/pyrimidine (PU) symmetry, designed for continuous-time Markov processes describing nucleotide substitution, particularly in phylogenetics. These models are characterized by their closure under matrix exponentiation and log-composition, their group-theoretic invariance under permutation of nucleotides preserving the purine/pyrimidine partition, and their suitability for modeling inhomogeneous evolutionary processes. The PU-Lie hierarchy systematically subsumes and extends classical models like Jukes–Cantor, Kimura 2- and 3-parameter, Felsenstein 81, and HKY within a unifying Lie algebraic framework (Fernández-Sánchez et al., 2012).

1. Continuous-Time Markov Chains and Model Motivations

Nucleotide substitution models in molecular evolution describe transitions across the four-state space X={A,G,C,T}\mathcal{X} = \{A,G,C,T\} using time-homogeneous Markov chains parameterized by rate matrices QQ. Constraints on QQ—for example, time-reversibility or rate symmetries—yield classical model families (JC, K80/K3ST, GTR, etc.). However, empirical data often require time-inhomogeneous rates, raising the problem of averaging Markov processes across evolutionary segments while preserving model class membership. This motivates the search for model classes that are closed under matrix logarithm of products—necessitating structure as Lie algebras—and admit biological symmetries such as purine–pyrimidine preservation (Fernández-Sánchez et al., 2012).

2. Lie Markov Models: Algebraic Structure and Closure

A Lie Markov model is a family of rate matrices LL such that for any stochastic Q1,Q2∈LQ_1,Q_2 \in L, log⁡(exp⁡(Q1)exp⁡(Q2))∈L\log(\exp(Q_1)\exp(Q_2)) \in L. The Baker–Campbell–Hausdorff formula implies LL must be a Lie subalgebra (under [Q1,Q2]=Q1Q2−Q2Q1[Q_1,Q_2] = Q_1 Q_2 - Q_2 Q_1) of the general Markov Lie algebra LGM\mathfrak{L}_{GM} (column-sum-zero matrices). Practical Lie Markov models additionally require existence of a basis with non-negative off-diagonal entries to retain the stochastic interpretation. Multiplicative closure ensures that time-varying (inhomogeneous) substitution processes can be represented as homogeneous ones within LL via the time-averaged log composition (Fernández-Sánchez et al., 2012).

3. Purine/Pyrimidine Symmetry and Model Constraints

Imposing PU symmetry requires invariance under permutations of nucleotides that preserve the partition QQ0 (purines) and QQ1 (pyrimidines). Specifically, the rate matrix QQ2 must satisfy linear constraints: rates for QQ3 (purine transitions), QQ4 (pyrimidine transitions), and various transversions are grouped according to the orbits of the PU-symmetry subgroup QQ5. This symmetry produces biologically interpretable models—parameters directly represent purine and pyrimidine transition rates and as well as various transversion rates (Fernández-Sánchez et al., 2012).

4. Decomposition and Classification of PU-Lie Models

The general Markov Lie algebra QQ6 decomposes under the action of the PU symmetry group into irreducible QQ7-modules, and every PU-Lie model is a QQ8-stable Lie subalgebra generated by a subset of these summands. The full classification enumerates these subalgebras by dimension and provides explicit parameterizations of the corresponding rate matrices. Notable classical models emerge as low-dimensional cases:

Dimension Model Name Distinctive Rate Structure
d=1 JC-like (Model 1.1) All off-diagonal rates equal
d=2 K80 (2.2b) Single purine, single transversion rate
d=3 K3ST (3.3a) Purine, pyrimidine, two transversion rates
d=4 F81 (4.4a) Four distinct rates by nucleotide
d=5,6,8+ HKY, K3ST+F81, … Increasingly refined symmetry patterns

Explicit forms for all PU-Lie models, including parameterizations and ray decompositions of the stochastic cone, are available in supplementary listings (Fernández-Sánchez et al., 2012).

5. Geometric Structure: The Stochastic Cone

Within the real form of the Markov Lie algebra, the subset of stochastic rate matrices forms a strongly convex polyhedral cone QQ9. The geometry of this cone (number and size of extremal rays, faces, orbits under PU symmetry) is crucial both for model interpretability (e.g., which evolutionary scenarios are captured) and for optimization in statistical inference pipelines. The dimension of QQ0 equals that of QQ1 precisely when a stochastic basis exists, and the invariance properties lead to systematic degenerations corresponding to model simplifications (Fernández-Sánchez et al., 2012).

6. Applications and Impact in Phylogenetic Inference

PU-Lie models support statistically coherent phylogenetic inference in scenarios with rate variation across evolutionary tree edges and time, since they ensure that the homogeneous “average” process of a complex history remains in the same class. The purine/pyrimidine symmetric hierarchy enables flexible selection of models matched to empirical substitution structure, and their linearity facilitates standard likelihood optimization. The symmetry structure allows for group-based Fourier methods and parameter identifiability, enhancing interpretability over generic GTR-like models (Fernández-Sánchez et al., 2012).

7. Broader Significance and Connections

The PU-Lie framework formalizes a principle: model families relevant to phylogenetics ought to form Lie algebras reflecting both their closure under time-inhomogeneous evolution and their biological symmetries. The explicit and exhaustive classification of PU-Lie models directly supersedes ad hoc parameterizations for purine/pyrimidine symmetric cases, embedding classical and new models within a single algebraic and geometric framework. This construction is broadly applicable to other group-based stochastic models requiring symmetry and closure properties, and suggests further extensions to amino-acid and codon evolution models using analogous symmetry analysis (Fernández-Sánchez et al., 2012).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to PU-Lie.