---
title: Reduced Factor Complexity Function
url: https://www.emergentmind.com/topics/reduced-factor-complexity-function
type: topic
---

# Reduced Factor Complexity Function

In combinatorics on words, the expression **reduced factor complexity function** is associated with several closely related reductions of the classical factor complexity of an infinite word. The most explicit formalization counts length-\(n\) factors only up to equality of their run-collapsed reductions \(\operatorname{red}(v)\). Other works use the phrase, or an interpretation in that spirit, for the increment \(C(n+1)-C(n)\), for symmetry quotients of factor sets by permutation groups, or for the reduction of possible recurrent spectra under linear factor-complexity bounds. This suggests that the topic is best understood as a family of constructions that simplify or constrain ordinary factor complexity while retaining combinatorial or dynamical structure [2509.16034] [0802.1332] [2607.07620] [2202.00643].

## 1. Run-based reduced factor complexity

Let \(w\) be a finite, nonempty word over a finite alphabet \(\mathcal{A}\). Writing \(w\) as a concatenation of maximal runs,
\[
w = c_1 \cdots c_1 \; c_2 \cdots c_2 \; \cdots \; c_n \cdots c_n,
\]
with \(c_i \neq c_{i+1}\), the reduction map is
\[
\operatorname{red}(w) = c_1 c_2 \cdots c_n.
\]
Thus \(\operatorname{red}(w)\) is obtained by replacing each maximal run by a single occurrence of its letter. Two finite words are reduced-equivalent if they have the same reduction:
\[
w \sim_{\mathrm{red}} v \quad\Longleftrightarrow\quad \operatorname{red}(w)=\operatorname{red}(v).
\]

For an infinite word \(\mathbf{x}=x_0x_1x_2\cdots\), let
\[
\mathcal{F}_n(\mathbf{x})=\{x_i x_{i+1}\cdots x_{i+n-1}: i\ge 0\}
\]
be the set of its length-\(n\) factors. The **reduced factor complexity** is
\[
\rho_{\mathbf{x}}^{\mathrm{red}}(n)
=
\left|
\left\{
\operatorname{red}(v): v\in \mathcal{F}_n(\mathbf{x})
\right\}
\right|.
\]
Equivalently, \(\rho_{\mathbf{x}}^{\mathrm{red}}(n)\) counts the \(\sim_{\mathrm{red}}\)-equivalence classes of length-\(n\) factors. The construction is a simplified version of ordinary factor complexity because it ignores multiplicities inside runs and retains only the pattern of letter-changes [2509.16034].

The same paper defines a reduced abelian analogue. Two length-\(n\) factors are equivalent for reduced abelian complexity when their reductions have the same Parikh vector, and the number of resulting classes is denoted \(\rho_{\mathbf{x}}^{\mathrm{ab},\mathrm{red}}(n)\). On a binary alphabet, \(\operatorname{red}(w)\) is always alternating, so a reduced-equivalence class is determined by the first letter and the number of runs. This binary specialization is central in the analysis of automatic examples [2509.16034].

## 2. Automatic-sequence case studies

For the Thue–Morse sequence \(\mathbf{t}\), the reduced factor complexity is significantly smaller and more regular than the classical factor complexity. The main recurrence proved is:
\[
\rho_{\mathbf{t}}^{\mathrm{red}}(n)
=
\begin{cases}
\rho_{\mathbf{t}}^{\mathrm{red}}\!\left(\dfrac{n+1}{2}\right),
& \text{if \(n\) is odd},\\[1ex]
\rho_{\mathbf{t}}^{\mathrm{red}}(m+1)+2,
& \text{if \(n=4m\) or \(n=4m+2\)}.
\end{cases}
\]
From this recurrence, \(\big(\rho_{\mathbf{t}}^{\mathrm{red}}(n)\big)_{n\in\mathbb{N}}\) is a \(2\)-regular sequence. The proof proceeds through alternation counts, with
\[
\#\text{runs}(w)=\#\text{alternations}(w)+1,
\]
and through recursive control of the minimum and maximum numbers of alternations among length-\(n\) factors [2509.16034].

For the regular paperfolding sequence \(\mathbf{f}\), the reduced factor complexity is bounded and eventually periodic. The explicit evaluation is
\[
\rho_{\mathbf{f}}^{\mathrm{red}}(n)
=
\begin{cases}
4 & \text{if } n \equiv 0,1,2,4,6 \pmod{8},\\
6 & \text{if } n \equiv 3,5,7 \pmod{8}.
\end{cases}
\]
The reduced abelian complexity of \(\mathbf{f}\) is likewise explicit:
\[
\rho_{\mathbf{f}}^{\mathrm{ab},\mathrm{red}}(n)
=
\begin{cases}
3 & \text{if \(n\) is even},\\
4 & \text{if \(n>1\) and \(n\equiv 1 \pmod{4}\)},\\
5 & \text{if \(n\equiv 3 \pmod{4}\)}.
\end{cases}
\]
Here the arguments combine the Toeplitz construction of \(\mathbf{f}\), run-count analysis, and Walnut to verify the existence of factors with the required run patterns for all \(n\) [2509.16034].

These examples show that run-based reduction can transform an unbounded linear complexity profile into either a \(2\)-regular sequence or a bounded periodic one. A plausible implication is that the reduction isolates a structural layer of the factor language that is more tightly governed by morphic or automatic self-similarity than the full factor set.

## 3. Reduced factor complexity as the increment \(C(n+1)-C(n)\)

A different usage identifies the **reduced factor complexity** with the first difference of the ordinary factor complexity,
\[
\Delta C(n)=C(n+1)-C(n),
\]
that is, the number of new factors of length \(n+1\) relative to length \(n\). In Rauzy-graph terms,
\[
C(n+1)-C(n)
=
\sum_{v\in S_n(w)} (\deg^+(v)-1),
\]
where \(S_n(w)\) is the set of special factors of length \(n\). This interpretation emphasizes branching rather than quotienting: the reduced function measures the excess of right extensions over the baseline value \(1\) [0802.1332].

For infinite words whose factor sets are closed under reversal, a fundamental equivalence relates this increment to palindromic structure. All complete returns to palindromic factors are palindromes if and only if
\[
P(n)+P(n+1)=C(n+1)-C(n)+2
\quad\text{for all }n,
\]
where \(P(n)\) denotes palindromic complexity. In this setting, the reduced factor complexity is exactly controlled by the palindromic contribution. Sturmian words satisfy this identity, and so do episturmian and Arnoux–Rauzy examples cited in that work [0802.1332].

A later asymptotic result gives a complementary comparison between palindromic and full factor complexity. For every non-ultimately periodic infinite word,
\[
\lim_{n\to\infty}
\frac{\Pal(n)\log(\Pal(n)+1)}{\rho(n)}=0,
\]
and the numerator is essentially optimal: for a nondecreasing \(f\), the universal implication
\[
\frac{f(\Pal(n))}{\rho(n)}\to 0
\]
for all non-ultimately periodic words holds if and only if \(f(t)=O(t\log(t+1))\) [2606.08127]. Taken together, these results separate two regimes. In reversal-closed rich words, \(C(n+1)-C(n)\) can be read directly from palindromic complexity, whereas in full generality palindromes remain asymptotically negligible compared with \(\rho(n)\).

## 4. Reduction by symmetry groups

Another formalization treats reduced factor complexity as a quotient of the factor set by a group action on positions. Let \(G_n \le S_n\) act on words \(a_1\cdots a_n\) by permuting coordinates, and let \(\sim_{G_n}\) be orbit equivalence. For a sequence \(\omega=(G_n)_{n\ge 1}\), the associated **group complexity** of an infinite word \(w\) is
\[
p_w^{\omega}(n)=|F_w(n)/{\sim_{G_n}}|.
\]
This construction interpolates between ordinary factor complexity and abelian complexity:
\[
p_w^{ab}(n)\le p_w^{G_n}(n)\le p_w(n),
\]
with \(G_n=Id_n\) giving factor complexity and \(G_n=S_n\) giving abelian complexity [2607.07620].

Within this framework, a word has **universal group complexity** if for every \(n\) and every integer \(m\) satisfying
\[
p_w^{ab}(n)\le m\le p_w(n),
\]
there exists a subgroup \(G\le S_n\) such that
\[
p_w^G(n)=m.
\]
Sturmian words satisfy this property. More precisely, if \(s\) is Sturmian, then for each \(n\) and each \(t\) with \(2\le t\le n+1\), there is a subgroup \(G\le S_n\) with \(p_s^G(n)=t\). The proof uses the subgroups \(S_{[m+1,n]}\) and the lexicographic array of factors [2607.07620].

The same paper studies aperiodic ternary words of minimal complexity \(p_x(n)=n+2\). Type I words have universal group complexity; Type II words fail universality at \(n=2\) but satisfy it for all \(n\ge 3\); and Type III words realize all intermediate values except possibly \(4\) when \(p_x^{ab}(n)=3\) [2607.07620]. This symmetry-based reduction is exact and tunable: it does not collapse runs or measure first differences, but instead identifies factors under prescribed positional symmetries.

## 5. Linear constraints, \(S\)-adic restrictions, and topological reduction

Several works interpret “reduction” not as a quotient on individual factors but as a global restriction on the possible complexity behavior of an infinite word. One setting is the topological invariant
\[
{\rm Rec}(w),
\]
constructed from recurrent right-infinite words whose factors are all factors of a given right-infinite word \(w\), modulo equality of factor sets. If \(w\) has linear factor complexity,
\[
p_w(n)\le Cn \quad\text{for all } n\ge 1,
\]
then \({\rm Rec}(w)\) is finite, and explicit upper bounds are proved:
\[
\#\operatorname{Rec}(w)
\le
\limsup_{n\to\infty}\big(p_w(n+1)-p_w(n)\big)+\lfloor C\rfloor!^2,
\]
\[
\#\operatorname{URec}(w)
\le
\limsup_{n\to\infty}\big(p_w(n+1)-p_w(n)\big)+1+C^2.
\]
The same work also shows that for every weakly increasing \(f:\mathbb{N}\to\mathbb{N}\) with \(f(n)\to\infty\), there exists a word \(w\) with \(p_w(n)=O(nf(n))\) and infinite \(\operatorname{Per}(w)\), hence infinite \({\rm Rec}(w)\). In that paper, the phrase “reduced factor complexity function” is not explicitly used, but the results are interpreted as showing that linear bounds reduce the recurrent spectrum to a finite topological space [2202.00643].

A related constrained setting appears for \(S\)-adic words generated by the Arnoux–Rauzy–Poincaré algorithm. There, unrestricted directive sequences in the substitution set can have quadratic complexity, whereas restricting to directive sequences accepted by the automaton \(G\) reduces the complexity to linear growth. For a totally irrational vector \(\mathbf{x}\in\Delta\), the associated \(S\)-adic word satisfies
\[
p(n+1)-p(n)\in\{2,3\}\quad\text{for all } n\ge 0,
\]
and
\[
2n+1\le p(n)\le \frac{5}{2}n+1 \quad\text{for all } n\ge 0.
\]
That paper explicitly notes that the phrase is not formally introduced as a new function; rather, the complexity is interpreted as reduced by the regular-language constraint on directive sequences [1404.4189].

These two developments make the same structural point in different languages. Linear constraints can force finiteness of a recurrent spectrum, while combinatorial constraints on admissible substitutions can force linear growth and tightly bounded increments. In both cases, the reduction acts on the space of admissible factor sets rather than solely on individual factors.

## 6. Open problems and current directions

The run-based theory leaves several questions open for the Thue–Morse sequence. The sequence \(\big(\rho_{\mathbf{t}}^{\mathrm{ab},\mathrm{red}}(n)\big)\) is listed explicitly in initial values, but a full recursion is unknown. It appears that
\[
\rho_{\mathbf{t}}^{\mathrm{ab},\mathrm{red}}(2n+1)
=
\rho_{\mathbf{t}}^{\mathrm{ab},\mathrm{red}}(n+1)
\quad\text{for all } n\ge 0,
\]
and a conjectured relation is proposed for
\[
\left|
\rho_{\mathbf{t}}^{\mathrm{ab},\mathrm{red}}(4n+2)
-
\rho_{\mathbf{t}}^{\mathrm{ab},\mathrm{red}}(4n)
\right|,
\]
but the sign of the difference is not known explicitly when nonzero. The paper also leaves open whether \(\big(\rho_{\mathbf{t}}^{\mathrm{ab},\mathrm{red}}(n)\big)\) is \(k\)-automatic for some base \(k\), and notes the suspicion that it is not \(k\)-automatic for any base [2509.16034].

The topological approach poses a realization problem: which finite topological spaces, equivalently finite posets via specialization order, can occur as \({\rm Rec}(w)\) for a word of linear factor complexity? The paper establishes constraints such as ACC, DCC, and bounds on minimal elements, but does not classify all realizable finite spaces [2202.00643].

The group-complexity framework leaves a specific gap for aperiodic ternary words of minimal complexity: in Type III, all intermediate values are realized except possibly \(4\) when \(p_x^{ab}(n)=3\). This unresolved value marks a precise obstruction to complete universality in the current classification [2607.07620].

Across these directions, the common theme is stable. A reduced factor complexity function is not a single invariant but a family of reductions of ordinary factor complexity: by collapsing runs, by taking first differences, by quotienting under group actions, or by imposing linear and \(S\)-adic constraints that drastically narrow the space of admissible factor behavior.

Source: https://www.emergentmind.com/topics/reduced-factor-complexity-function