---
title: Strand Symmetric Model in Phylogenetics
url: https://www.emergentmind.com/topics/strand-symmetric-model
type: topic
---

# Strand Symmetric Model in Phylogenetics

The strand symmetric model is a class of DNA sequence evolution models defined by invariance under Watson–Crick complementation, \(A \leftrightarrow T\) and \(C \leftrightarrow G\). In phylogenetics, it is a \(4\)-state Markov model in which substitution parameters related by complementation are identified, placing it between the fully general \(4\)-state Markov model and more restrictive models such as Jukes–Cantor or Kimura-type models. In algebraic statistics, its Zariski closure has a particularly structured description, and in continuous-time formulations it forms a multiplicatively closed model with a six-parameter rate matrix. Related work also uses “strand symmetry” more broadly for sequence-generation models designed to explain intra-strand parity and reverse-complement regularities in bacterial genomes [1410.3344].

## 1. Biological basis and core symmetry

The defining biological premise is that DNA is double-stranded and that substitutions related by exchanging each nucleotide with its complement should have equal probabilities. In the phylogenetic formulation, this means that if \(\overline{A}=T\), \(\overline{T}=A\), \(\overline{C}=G\), and \(\overline{G}=C\), then strand symmetry requires
\[
\theta^{(e)}_{ij}=\theta^{(e)}_{\overline{i}\,\overline{j}}
\]
for each edge \(e\) and each pair of nucleotides \(i,j\). The root distribution is likewise constrained by
\[
\pi_A=\pi_T,\qquad \pi_C=\pi_G.
\]
These equalities encode Watson–Crick symmetry without imposing time-reversibility or equal base frequencies beyond the complement constraints [1410.3344].

In explicit matrix form, one common discrete-time parameterization writes a strand-symmetric Markov matrix as
\[
M=\begin{pmatrix}
a & b & c & d\\
e & f & g & h\\
h & g & f & e\\
d & c & b & a
\end{pmatrix},\qquad a=1-b-c-d,\qquad f=1-e-g-h.
\]
The corresponding continuous-time rate matrix can be written with six free off-diagonal parameters. In one ordering, \((A,T,G,C)\), this takes the form
\[
Q=
\begin{pmatrix}
-(a+c+e) & a & c & e\\
a & -(a+c+e) & e & c\\
b & d & -(b+d+f) & f\\
d & b & f & -(b+d+f)
\end{pmatrix},
\]
showing directly that the \(12\) off-diagonal entries of a general \(4\times 4\) rate matrix are reduced to \(6\) under strand symmetry [1905.02015].

A recurrent point in the literature is that strand symmetry is distinct from reversibility. The model may be non-reversible, and detailed balance is not implied by the complement equalities alone. This is a central conceptual distinction from GTR: the strand symmetric model has six rate parameters like GTR, but its constraints are organized by complementarity rather than exchangeability plus a stationary distribution [1905.02015].

## 2. Statistical formulation on phylogenetic trees

For a rooted phylogenetic tree \(T\) with node set \(V\) and edge set \(E\), each node \(v\in V\) carries a random variable \(X_v\) taking values in \(\Sigma=\{A,C,G,T\}\). The root has distribution \(\pi\), and each directed edge \(e=(u\to v)\) has a transition matrix \(M^{(e)}=(\theta_{ij}^{(e)})\), with row sums equal to one and with the strand-symmetry constraints described above. For an \(n\)-leaf tree and a leaf pattern \(\mathbf{i}=(i_1,\dots,i_n)\in\Sigma^n\), the joint distribution is given by the usual Markov-tree sum over hidden ancestral states:
\[
p_{\mathbf{i}}=\sum_{(x_v)_{v\in V}:\,x_{\text{leaves}}=\mathbf{i}}
\pi_{x_r}\prod_{e=(u\to v)\in E}\theta^{(e)}_{x_u x_v}.
\]
Collecting all \(p_{\mathbf{i}}\) yields a point in \(\mathbb{R}^{4^n}\), or in the simplex after imposing stochastic normalization [1410.3344].

In this setting, the strand symmetric model is a restriction of the general time-inhomogeneous \(4\times 4\) Markov model. Each edge still has its own transition matrix, but entries paired by complement are identified. The model is therefore less restrictive than Jukes–Cantor and related classical symmetric models, because it does not require time-reversibility, yet it is substantially more structured than the unrestricted general Markov model [1410.3344].

For the three-leaf claw tree \(K_{1,3}\), which is the fundamental local object in many algebraic-statistical analyses, the model map specializes to
\[
\phi_{K_{1,3}}:S_{K_{1,3}}\to\mathbb{R}^{4^3},\qquad p_{ijk}=\mathbb{P}(X_1=i,X_2=j,X_3=k).
\]
This tree is pivotal because general results for matrix-valued group-based models imply that, once the defining equations are known for the claw tree, one can determine the vanishing ideals for arbitrary trivalent trees by gluing and pullback constructions [1410.3344].

A different but closely related phylogenetic formulation studies non-stationary strand-symmetric models on rooted trees. There, each internal node has its own marginal distribution, and the root distribution is not required to be stationary for any edge process. In that framework, the strand-symmetry constraint is often written compactly as
\[
P^{uv}=SP^{uv}S\qquad\text{or}\qquad Q^{uv}=SQ^{uv}S,
\]
where
\[
S=
\begin{pmatrix}
0&0&0&1\\
0&0&1&0\\
0&1&0&0\\
1&0&0&0
\end{pmatrix}.
\]
This matrix implements reversal of nucleotide order so that the symmetry becomes conjugation by \(S\) [1610.05057].

## 3. Algebraic-geometric structure and defining equations

A major algebraic-statistical result is that, for the strand symmetric model on the three-leaf claw tree, the previously known phylogenetic invariants completely define the vanishing ideal. In Fourier coordinates arising from a matrix-valued group-based realization over \(G=\mathbb{Z}_2\), the parameterization becomes sparse:
\[
q^{mno}_{ijk}=
\begin{cases}
d^{mm}_{0i}e^{nn}_{0j}f^{oo}_{0k}+d^{mm}_{1i}e^{nn}_{1j}f^{oo}_{1k} & \text{if } m+n+o\equiv 0\pmod 2,\\[4pt]
0 & \text{otherwise}.
\end{cases}
\]
Only coordinates with \(m+n+o\equiv 0\) can be nonzero, and each nonzero coordinate is a sum of two monomials. This identifies the model variety as a secant variety of a toric variety [1410.3344].

More precisely, if \(C\) denotes the toric variety obtained by suppressing one of the two summands, then the affine cone over the projective model satisfies
\[
CV(T)=C*C.
\]
The toric variety \(C\) is parameterized by monomials of the form
\[
d^{mm}_{0i}e^{nn}_{0j}f^{oo}_{0k},
\]
and can be viewed as arising from a Segre embedding
\[
\mathbb{P}^3\times\mathbb{P}^3\times\mathbb{P}^3\hookrightarrow\mathbb{P}^{63}
\]
followed by coordinate projection determined by the parity restriction \(m+n+o\equiv 0\) [1410.3344].

The central theorem establishes that the ideal generated by the known \(32\) degree-three polynomials and \(18\) degree-four polynomials is the full vanishing ideal for the claw tree. If \(I\) denotes the ideal generated by these \(50\) invariants, then
\[
\dim I=20,\qquad \deg I=9024,
\]
and the Hilbert series is
\[
\frac{1 + 12t + 78t^2 + 332t^3 + 984t^4 + 1908t^5 + 2394t^6 + 2394t^7 + 1908t^8 + 984t^9 + 332t^{10} + 78t^{11} + 12t^{12} + t^{13}}{(1-t)^{20}}.
\]
The proof combines a secant-variety dimension argument with elimination theory and a computational primality proof. A \(12\times 32\) exponent matrix \(A\) of rank \(10\) is used with Draisma’s tropical secant dimension technique, yielding the lower bound \(\dim(C*C)\ge 19\) in projective space and hence affine dimension at least \(20\). Since the variety cut out by the \(50\) known equations also has dimension \(20\), and the corresponding ideal is shown to be prime, the two coincide [1410.3344].

This result has a structural consequence beyond the three-leaf case. Because the strand symmetric model fits the matrix-valued group-based framework, knowledge of the claw-tree ideal determines the vanishing ideal for any trivalent tree through edge-splitting and gluing operations. The claw tree is therefore not merely a minimal example but the local generator of the algebraic theory for the entire class [1410.3344].

## 4. Matrix group structure and Markov invariants

The continuous-time strand symmetric phylogenetic substitution model also has a Lie-theoretic description. Its rate matrices form a six-dimensional Lie algebra generated by matrices \(S_1,S_2,S_3,T_1,T_2,T_3\), and the corresponding set of Markov matrices is multiplicatively closed. This permits a decomposition of the Lie algebra as
\[
l_{SSM}\cong sl_2\oplus gl_1\oplus l_2,
\]
where \(l_2\) is the two-dimensional nonabelian shift algebra [1307.5574].

A particularly important change of basis is the “split” basis
\[
u_0=e_A-e_T,\qquad u_1=e_C-e_G,\qquad
v_{0'}=e_A+e_T,\qquad v_{1'}=e_C+e_G.
\]
In this basis, the four-dimensional nucleotide state space decomposes as
\[
\mathcal{S}=U\oplus V,
\]
with \(U=\langle u_0,u_1\rangle\) and \(V=\langle v_{0'},v_{1'}\rangle\). Strand-symmetric Markov matrices then take the block form
\[
M=m\oplus\overline{m}=
\begin{pmatrix}
m&0\\
0&\overline{m}
\end{pmatrix}.
\]
This exhibits the model as a direct sum of two \(2\times 2\) components, one acting on difference coordinates and one on sum coordinates [1307.5574].

The block decomposition leads to a systematic classification of Markov invariants. For \(L\) taxa and polynomial degree \(D\), representation-theoretic methods using plethysm and symmetric-group inner products enumerate all one-dimensional subrepresentations. In the quadratic and cubic cases, the model admits exactly
\[
\frac{1}{3}(3^L+(-1)^L)
\]
and
\[
6^{L-1}
\]
linearly independent Markov invariants, respectively [1307.5574].

For \(L=2\), the nontrivial quadratic invariants are the determinants of the four \(2\times 2\) blocks of the \(4\times 4\) pattern matrix in split coordinates:
\[
f^{(1,1)}=P_{00}P_{11}-P_{01}P_{10},
\]
\[
f^{(1,2)}=P_{00'}P_{11'}-P_{01'}P_{10'},
\]
\[
f^{(2,1)}=P_{0'0}P_{1'1}-P_{0'1}P_{1'0},
\]
\[
f^{(2,2)}=P_{0'0'}P_{1'1'}-P_{0'1'}P_{1'0'}.
\]
Under leafwise action, these transform by products of \(\det(m_i)\) and \(\det(\overline{m}_i)\), so they are genuine Markov invariants rather than vanishing invariants [1307.5574].

The same paper relates the block determinants to two biologically interpretable rate aggregates. If
\[
\sigma_1=\alpha_3+\beta_3
\quad\text{and}\quad
\sigma_2=\alpha_1+\alpha_2+\beta_1+\beta_2,
\]
then for a two-taxon tree
\[
\det(m)=e^{-(2\sigma_1+\sigma_2)t},\qquad
\det(\overline{m})=e^{-\sigma_2 t}.
\]
The quadratic invariants therefore provide independent estimates of phylogenetic distances associated with substitution rates within Watson–Crick conjugate pairs and across conjugate base pairs [1307.5574].

## 5. Non-stationarity, identifiability, and rooting

A distinct line of work analyzes non-stationary strand-symmetric models on rooted phylogenies and proves identifiability of both the root location and the full model. The crucial non-stationarity assumption is that internal-node marginals are compositionally asymmetric. In one formalization, a probability vector \(\pi\) is compositionally asymmetric if all entries are non-zero and
\[
\pi(1)\neq\pi(4),\qquad
\pi(2)\neq\pi(3),\qquad
\pi(1)/\pi(4)\neq\pi(2)/\pi(3),\qquad
\pi(1)/\pi(4)\neq\pi(3)/\pi(2).
\]
A stationary strand-symmetric distribution violates these constraints, so the condition explicitly excludes the stationary case [1610.05057].

For the rooted two-taxon tree with root \(m\) and children \(a,b\), the joint tip distribution is
\[
J^{ab}=(P^{mb})^\top \Pi^m P^{ma},\qquad \Pi^m=\mathrm{diag}(\pi^m).
\]
Using the permutation matrix \(S\), one constructs
\[
G=S(J^{ab})^{-1}SJ^{ab},
\]
and obtains the factorization
\[
G=(P^{ma})^{-1}
\mathrm{diag}\!\left(
\frac{\pi^m(1)}{\pi^m(4)},
\frac{\pi^m(2)}{\pi^m(3)},
\frac{\pi^m(3)}{\pi^m(2)},
\frac{\pi^m(4)}{\pi^m(1)}
\right)
P^{ma}.
\]
Because compositionally asymmetric root distributions give distinct eigenvalues, this eigendecomposition recovers \(\pi^m\), \(P^{ma}\), and \(P^{mb}\) up to the unavoidable internal-state relabellings [1610.05057].

This yields an unusual identifiability result: under strand symmetry, non-stationarity, invertibility, and non-permutation assumptions, two sequences are already sufficient to identify the full rooted two-taxon model. More generally, on arbitrary rooted trees, the unrooted topology is reconstructed from the additive “paralinear distance”
\[
f(\{r,s\})=-\log\det(P^{rs}P^{sr}),
\]
and the root is then located using pairwise distributions and most-recent-common-ancestor arguments. The full model is identifiable provided edge matrices lie in a class reconstructible from rows, together with the more general notion of “sympathetic” matrices that support consistent internal-state labelling without requiring every edge to satisfy a strict diagonal-largest-in-column condition [1610.05057].

The same framework establishes statistical consistency of maximum likelihood. Under strong reconstructibility and strong sympathy conditions, the maximum-likelihood estimators of the rooted topology, root marginal distribution, and edge transition matrices converge to their true values as alignment length tends to infinity. A continuous-time version follows when each edge transition matrix has a unique logarithm with non-negative off-diagonals and strand-symmetric rate matrix [1610.05057].

A common misconception is that non-reversibility alone suffices to root trees. The non-stationary strand-symmetric theory explicitly rejects that conclusion: stationary but non-reversible strand-symmetric models can still be non-identifiable for root placement, whereas non-stationarity together with strand symmetry breaks the usual rooting ambiguity [1610.05057].

## 6. Mutation–drift equilibrium and parameter inference

In population-genetic form, the strand symmetric model is studied as a mutation matrix in a neutral mutation–drift equilibrium with small scaled mutation rates. Under the six-parameter rate matrix
\[
Q=
\begin{pmatrix}
-(a+c+e) & a & c & e\\
a & -(a+c+e) & e & c\\
b & d & -(b+d+f) & f\\
d & b & f & -(b+d+f)
\end{pmatrix},
\]
the stationary distribution is
\[
\pi_A=\pi_T=\frac{b+d}{2(b+c+d+e)},\qquad
\pi_G=\pi_C=\frac{c+e}{2(b+c+d+e)}.
\]
Thus Chargaff’s second parity rule holds at stationarity for this mutation model [1905.02015].

Assuming small scaled mutation rates, the stationary sampling distribution for a sample of size \(M\) can be expanded to first order in the mutation parameter. For a general \(K\)-allele model, monomorphic and biallelic sample configurations dominate at \(O(\theta)\), while configurations with three or more alleles are \(O(\theta^2)\). In the strand-symmetric case this yields explicit first-order probabilities for monomorphic \(A\), \(T\), \(G\), and \(C\) sites and for each of the six biallelic classes \(A/T\), \(A/C\), \(A/G\), \(T/C\), \(T/G\), and \(G/C\) [1905.02015].

The estimation strategy proceeds in stages. First, nucleotides are grouped into compound alleles \(AT\) and \(GC\), giving a biallelic model with
\[
\beta=\frac{b+d}{b+d+c+e},\qquad
1-\beta=\frac{c+e}{b+d+c+e},
\]
and
\[
\gamma=\beta(1-\beta)\theta=\frac{(b+d)(c+e)}{b+d+c+e}.
\]
If \(L_0\), \(L_M\), and \(L_p\) are the counts of \(GC\)-fixed, \(AT\)-fixed, and polymorphic \(AT/GC\) sites, then the first-stage maximum-likelihood estimators are
\[
\widehat{\beta}=\frac{L_M+L_p/2}{L_0+L_M+L_p},\qquad
\widehat{\gamma}=\frac{L_p}{2(L_0+L_M+L_p)H_{M-1}},
\]
where \(H_{M-1}\) is the harmonic number [1905.02015].

Second, the within-pair symmetric rates are estimated from the \(A/T\) and \(G/C\) subsystems:
\[
\hat{a}=
\frac{L_p^{(AT)}L_M^{(AT)}}
{\big(L_M^{(AT)}+L_p^{(AT)}+L_0^{(AT)}\big)\big(L_M^{(AT)}+L_p^{(AT)}/2\big)H_{M-1}},
\]
\[
\hat{f}=
\frac{L_p^{(GC)}L_0^{(GC)}}
{\big(L_M^{(GC)}+L_p^{(GC)}+L_0^{(GC)}\big)\big(L_0^{(GC)}+L_p^{(GC)}/2\big)H_{M-1}}.
\]
Third, the cross-pair parameters \((e,d)\) and \((c,b)\) are estimated by EM algorithms derived from the \(A/C\) and \(A/G\) biallelic subsystems, respectively. This is presented as the first time that ML estimators are provided for a mutation model more complex than parent-independent mutation [1905.02015].

The population-genetic strand symmetric model is therefore not merely a formal analogue of the phylogenetic model. It provides an explicit inferential framework in which the six complement-constrained mutation parameters can be estimated from the site frequency spectrum under mutation–drift equilibrium [1905.02015].

## 7. Related sequence-construction models and broader usage

Outside phylogenetics, “strand symmetry” also names models intended to explain intra-strand parity and reverse-complement regularities in whole genomes. A prominent example is the Sobottka–Hart construction for primitive double-stranded DNA. There the model is defined by a strand-symmetric probability vector \(M\) and a persymmetric acceptance matrix \(L\),
\[
M_j=M_{\alpha(j)},\qquad
L_{ij}=L_{\alpha(j)\alpha(i)},
\]
with \(\alpha(A)=T\), \(\alpha(T)=A\), \(\alpha(C)=G\), and \(\alpha(G)=C\). From these, a Markov transition matrix
\[
P_{ij}=\frac{L_{ij}M_j}{\sum_k L_{ik}M_k}
\]
is induced on each semi-strand, while a whole strand is modeled as a concatenation of a process and its reverse complement with mixing weight \(t\approx 1/2\) [2206.00610].

In the related hidden-Markov construction for bacterial DNA, the same logic appears in terms of an environmental availability vector \(\mu\) and an acceptance matrix \(\aleph\), with
\[
\mu_A=\mu_T,\qquad \mu_C=\mu_G,
\]
and
\[
a_{ij}=a_{\alpha(j)\alpha(i)}.
\]
The effective transition matrix along a constructed segment is
\[
W_{ij}=\frac{a_{ij}\mu_j}{\sum_k a_{ik}\mu_k},
\]
and the global mono- and dinucleotide frequencies of a strand are given by the mixture formulas
\[
P_{ij}=tQ_{ij}+(1-t)Q_{\alpha(j)\alpha(i)},\qquad
\pi_i=t\nu_i+(1-t)\nu_{\alpha(i)},
\]
with complementary expressions for the opposite strand. When \(t=1/2\), the model gives \(P=R\) and \(\pi=\rho\), so Chargaff’s second parity rule for monomers and dimers follows at the level of observed statistics [1401.3254].

These constructions differ from classical phylogenetic strand symmetric models in a precise way. They are not substitution models on trees and do not assume a single homogeneous process across the entire sequence. Instead, they model a sequence as a concatenation of two complementary Markov segments. This yields whole-strand intra-strand parity together with half-strand asymmetry, which the 2022 analysis presents as a joint explanation of intra-strand parity and strand compositional asymmetry in bacterial genomes [2206.00610].

A plausible implication is that the phrase “strand symmetric model” has two technically related but non-identical usages. In phylogenetics and algebraic statistics it denotes a complement-constrained substitution model on trees; in bacterial genome composition studies it denotes a complement-constrained sequence-generation or sequence-growth mechanism. The shared mathematical core is the same persymmetry or reverse-complement symmetry, but the modeled objects differ: substitution along a phylogeny in one case, and compositional structure along a genome coordinate in the other [1401.3254].

In the phylogenetic literature, the strand symmetric model is therefore best understood as a non-reversible, complement-constrained \(4\)-state Markov model whose algebraic geometry, Lie structure, identifiability theory, and inferential procedures are unusually explicit. In the broader DNA-sequence literature, the same symmetry principle underlies models of intra-strand parity and reverse-complement regularity, reinforcing the view that Watson–Crick complementarity can be encoded as a mathematically precise invariance principle across several distinct modeling regimes [1307.5574].

Source: https://www.emergentmind.com/topics/strand-symmetric-model