---
title: Inharmonic Series Model in Signal Analysis
url: https://www.emergentmind.com/topics/inharmonic-series-model
type: topic
---

# Inharmonic Series Model in Signal Analysis

An inharmonic series model is a representation of a signal or dynamical system in which prominent components are not constrained to lie exactly at integer multiples of a single fundamental. In signal analysis, the canonical almost-harmonic formulation is
$$
x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad \omega_k=k\omega_0+\Delta_k,
$$
where the $\Delta_k$ quantify departures from exact harmonicity. The literature does not treat this as a single universal construction. Instead, it includes deterministic perturbations of a harmonic lattice, definitions of a nominal fundamental based on waveform fit or spectral transport, stochastic perturbation models, and phase–waveform models that avoid explicit partial decomposition altogether [2003.10767; 0911.5171].

## 1. Canonical almost-harmonic formulation

The basic signal-processing form of an inharmonic series model writes a finite-length or discrete-time signal as a sum of complex exponentials with frequencies near a harmonic scaffold,
$$
x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad t=0,1,\ldots,N-1,
$$
with
$$
\omega_k=k\omega_0+\Delta_k.
$$
In the perfectly harmonic case, all $\Delta_k$ vanish and $\omega_0$ is an ordinary fundamental frequency. In the inharmonic case, the $\Delta_k$ are assumed small in the sense that $\Delta_k\ll\omega_0$. This model is the standard “almost harmonic” representation: each partial is a perturbation of an integer multiple of a nominal base frequency rather than an unrelated sinusoid [2003.10767].

The same framework admits structured special cases. A stiff-string example is
$$
\omega_k=k\omega_0\sqrt{1+k^2\beta},
$$
so that
$$
\Delta_k=k\omega_0\left(\sqrt{1+k^2\beta}-1\right)>0.
$$
Here the inharmonicity is systematic rather than arbitrary. A generic harmonic approximant used for analysis is
$$
\mu_t=\sum_{\ell=1}^L r_\ell e^{i\phi_\ell+i\omega_0\ell t},
$$
with $L$ not necessarily equal to $K$. This approximant is central because many estimation procedures operate on the harmonic model even when the true signal is only almost harmonic [2003.10767].

A plausible implication is that the phrase “inharmonic series model” is best understood as a family of nearby constructions centered on a nominal harmonic lattice, not as a single parametrization. Some models treat the deviations $\Delta_k$ as deterministic offsets, others as random variables, and others replace explicit partials by a different state space entirely.

## 2. Fundamental frequency in inharmonic signals

A central difficulty is that an inharmonic signal has no unique obvious notion of fundamental frequency. One formalization defines $\omega_0$ as the best $\ell_2$ harmonic approximation:
$$
\omega_0=\arg\min_{\theta}\frac{1}{N}\sum_{t=0}^{N-1}|x_t-\mu_t(\theta)|^2.
$$
Under a misspecified Gaussian harmonic model, this is exactly the pseudo-true parameter, namely the Kullback–Leibler best approximation inside the harmonic family. In that setting, the pseudo-true noise variance becomes
$$
\tilde{\sigma}^2=\sigma^2+\frac{1}{N}\sum_{t=0}^{N-1}|\epsilon_t(\theta_0)|^2.
$$
This definition is useful as a benchmark for harmonic estimators applied to inharmonic data, but it depends on observation length and nuisance phases, and it can become ambiguous as $N\to\infty$ because different subharmonic candidates may fit equally well [2003.10767].

A second definition replaces waveform fit by spectral transport. With random initial phases so that the process is wide-sense stationary, the inharmonic spectrum is
$$
\Phi_x(\omega)=2\pi\sum_{k=1}^K r_k^2\delta(\omega-\omega_k),
$$
and a harmonic spectrum is
$$
\Phi_\mu(\omega)=2\pi\sum_{\ell=1}^L r_\ell^2\delta(\omega-\ell\omega_0).
$$
The distance is defined by Monge–Kantorovich transport with cost
$$
c(\omega_1,\omega_2)=(\omega_1-\omega_2)^2.
$$
For fixed $\omega_0$, the optimization reduces to
$$
\min_{\omega_0}2\pi\sum_{k=1}^K r_k^2\min_{\ell\in\{1,\ldots,L\}}(\ell\omega_0-\omega_k)^2.
$$
For sufficiently small perturbations, the minimizing $\omega_0$ is explicitly
$$
\omega_0=\frac{\sum_{k=1}^K r_k^2 k\omega_k}{\sum_{k=1}^K r_k^2 k^2}
=\tilde{\omega}_0+\frac{\sum_{k=1}^K r_k^2 k\Delta_k}{\sum_{k=1}^K r_k^2 k^2}.
$$
This is phase-free and observation-length-free, and it interprets the nominal fundamental as a power-weighted projection of perturbed partial frequencies onto the harmonic manifold [2003.10767].

A third definition treats the inharmonic series as a randomly perturbed harmonic signal,
$$
x_t=\sum_{k=1}^K r_k e^{i\phi_k+i(k\omega_0+\Delta_k)t},
$$
with independent zero-mean perturbations, often specialized to
$$
\Delta_k\sim\mathcal N(0,\sigma_\Delta^2).
$$
Here $\omega_0$ is the mean harmonic scaffold, and the appropriate lower bound is a hybrid Cramér–Rao lower bound because the model mixes deterministic parameters with random latent perturbations. The associated ML/MAP criterion combines an unstructured sinusoidal data fit with a penalty on departures from a harmonic lattice. This yields a bridge between perfectly harmonic and fully unstructured sinusoidal models [2003.10767].

## 3. Waveform-field and phase–time models

An alternative tradition does not represent a monophonic sound as a sum of partials at all. Instead, it models the signal as a function on a cylinder with independent shape-time and phase coordinates,
$$
y\in\mathcal F(\mathbb R\times\mathbb T,V),
$$
where $\mathbb T=\mathbb R/\mathbb Z$. Ordinary audio is obtained by sampling the cylindrical field along a control path,
$$
S_{h,g}y(t)=y(h(t),g(t)).
$$
Here $h(t)$ controls shape or timbral progression, while $g(t)$ controls phase progression. The source signal is lifted to the cylinder by interpolation,
$$
Fx(t,\varphi)=\sum_{\tau\in\varphi}x(\tau)\,\kappa(t-\tau).
$$
This construction is explicitly presented as robust against non-harmonic frequencies because it does not require decomposition into harmonically locked sinusoids [0911.5171].

For a sinusoidal input $x(t)=e^{2\pi iat}$, the model decomposes the frequency as
$$
a=b+n,\qquad n\in\mathbb Z,\quad b\in\left(-\tfrac12,\tfrac12\right),
$$
and, with Whittaker interpolation, the resynthesized frequency under shape-speed $v$ and phase-speed $\alpha$ is
$$
a\mapsto bv+n\alpha.
$$
This preserves pure sines and gives a predictable mapping for non-integer frequencies without requiring a fixed harmonic series. In this representation, the integer part of a component is tied to phase progression and the residual detuning to shape-time progression. The model therefore encodes inharmonicity as residual detuning or waveform evolution across the shape axis rather than as an explicit bank of inharmonic partial tracks [0911.5171].

The cylindrical framework also proves exact structural properties. It is linear, it preserves a static periodic wave as a constant waveform around the cylinder, and it contains ordinary resampling as a special case. If the input has the separable form “envelope times waveform,” then the model preserves that factorization:
$$
F(f\cdot (w\circ c))(t,\varphi)=f(t)\,w(\varphi),
$$
hence
$$
S_{h,g}(F(f\cdot (w\circ c)))=(f\circ h)\cdot (w\circ g).
$$
Its stated limitations are equally clear: it is aimed at monophonic signals with a known constant period, it excludes noise portions such as speech in the exact theory, it is not polyphonic, and it does not preserve formant structure [0911.5171].

## 4. Estimation and inverse modeling

One important use of an inharmonic series model is not synthesis but estimation under model mismatch. A misspecified harmonic estimator assumes
$$
y_t=\mu_t(\theta)+w_t=\sum_{k=1}^K \rho_k e^{i\varphi_k+ik\omega_0 t}+w_t,
$$
even when the true data come from an inharmonic process
$$
y_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t}+e_t,\qquad \omega_k=k\omega+\Delta_k.
$$
In the Gaussian setting, the pseudo-true parameter is the least-squares harmonic projection,
$$
\theta_0=\arg\min_\theta \sum_{t=0}^{N-1}|x_t-\mu_t(\theta)|^2,
$$
and the pseudo-true variance is inflated by deterministic mismatch:
$$
\tilde{\sigma}_0^2=\sigma^2+\frac{1}{N}\sum_{t=0}^{N-1}|\xi_t(\theta_0)|^2.
$$
The misspecified Gaussian Cramér–Rao lower bound then quantifies the bias–variance tradeoff of estimating an inharmonic signal through a harmonic scaffold. The stated conclusion is that, for moderately inharmonic signals, the misspecified harmonic model can achieve lower mean squared error than a correctly specified but unstructured sinusoidal model, and voiced speech is reported to belong to this class [1910.07016].

A more recent inverse formulation keeps the signal approximately harmonic,
$$
\omega_\ell^{(k)}=\ell\omega_0^{(k)}+\Delta_\ell^{(k)},\qquad |\Delta_\ell^{(k)}|\ll \omega_0^{(k)},
$$
but does not fit a parametric law for the $\Delta_\ell^{(k)}$. Instead, it represents spectral content as a measure on the unit circle and assigns spectral mass to candidate fundamentals through optimal transport. The normalized ground cost is
$$
c(\omega,\omega_0)=\min_{k\in\mathbb Z_+}\left(\frac{\omega}{\omega_0}-k\right)^2.
$$
Group sparsity over candidate fundamentals is enforced with
$$
\|m\|_{\infty,1}=\sum_{g=1}^G \|P^{(g)}(m)\|_\infty.
$$
This decouples the measurement operator from the harmonic prior: the forward model uses Fourier atoms, while harmonicity is enforced in the regularizer. The stated consequence is robustness to inharmonicity without introducing a stretched-comb law such as $f_k=kf_0\sqrt{1+Bk^2}$ [2508.02471].

These estimation frameworks share a common geometry. The signal is not forced to be exactly harmonic; rather, it is projected, penalized, or softly clustered relative to a harmonic scaffold. This suggests that many practical “inharmonic series models” are best read as structured approximations, not as claims of exact periodicity.

## 5. Audio, speech, and perceptual consequences

The computational importance of harmonic versus inharmonic structure is especially visible in monaural speech segregation. One study defines harmonic speech as having Fourier components at the fundamental frequency $F0$ and its integer multiples, and generates inharmonic speech by jittering each harmonic:
$$
f_n(t)=n f_0(t)+J_n f_0(t),\qquad -J\le J_n\le J.
$$
The tested range is
$$
J\in[0.01,0.30].
$$
Under this slight departure from exact harmonicity, end-to-end separators degrade sharply: Conv-TasNet is reported to drop from $15.4$ dB to $0.70$ dB under $3\%$ harmonic jitter, and DPT-Net from $20.2$ dB to $0.7$ dB for imperceptible amounts of harmonic jitter. Training on inharmonic speech does not remove the sensitivity and worsens natural-speech performance. The paper interprets this as evidence that these networks rely strongly on harmonic spectral patterns, whereas temporal-coherence systems rely primarily on timing cues and are much less tied to exact harmonic spacing [2203.04420].

Inharmonic series models also matter for roughness and consonance calculations. A Helmholtz/Plomp–Levelt dissonance formulation operates directly on arbitrary partial frequencies and amplitudes. For a pair of partials,
$$
x=\frac{|f_1-f_2|}{b(f)},
$$
and the dyadic roughness is
$$
g(x)=4.906\,x(1.2-x)^4,\qquad 0\le x<1.2,
$$
with $g(x)=0$ for $x\ge 1.2$. Total dissonance is not obtained by naive summation but by a loudness/intensity-weighted aggregate, with preferred exponent $\alpha=\tfrac12$. Because the formulas are expressed in terms of arbitrary partial frequencies $f_i,f_j$ and intensities, the method can be applied directly to inharmonic spectra rather than only to exact harmonic tones [1306.1344].

Together these results indicate two distinct uses of the concept. In one use, inharmonicity is a nuisance that destabilizes systems trained on exact harmonic structure. In the other, it is the primary object of analysis, as in roughness computation or robust pitch estimation.

## 6. Broader physical and mathematical extensions

Outside signal processing, related models appear in nonlinear dynamics, mathematical physics, and auditory modeling. In a delay-differential toy model of flute-like instruments, the resonator is written as
$$
Y(\omega)=\sum_{k=1}^m \frac{a_k\,j\omega}{\omega_k^2-\omega^2+j\omega\,\omega_k/Q_k},
$$
and inharmonicity is controlled by the ratio $\omega_2/\omega_1$. The cases $\omega_2/\omega_1=1.99$, $2$, and $2.05$ produce qualitatively different behavior: bistable hysteresis, an intermediate quasiperiodic regime, or first-register dominance. Here an “inharmonic series model” means a low-order modal resonance series whose detuning actively changes nonlinear regime selection [1207.7136].

For a periodically forced anharmonic chain, the asymptotic periodic state is represented as a convergent power series in the anharmonicity parameter,
$$
q_{x,\mathrm p}(t;\nu)=\sum_{l=0}^{\infty} q_x^{(l)}(t;\nu)\,\nu^l.
$$
The expansion is proved to converge for $|\nu|<\nu_0$ under the nonresonance condition that no integer multiple of the forcing frequency falls into the harmonic spectral band
$$
[\omega_0,\sqrt{\omega_0^2+4}].
$$
In this setting, “anharmonic series model” refers not to detuned partials but to a perturbation series for the unique periodic solution of a nonlinear lattice [2503.23527].

A one-dimensional polaron model on an inharmonic lattice uses the cubic FPU-type potential
$$
U(q)=\frac{q^2}{2}-\frac{\beta q^3}{3}.
$$
Its continuum equations admit exact one-soliton solutions only on a special relation between the electron–phonon and inharmonicity parameters, while numerical simulations in a stronger inharmonic regime produce unusual stable moving polarons with multi-peak envelopes and, for $(\alpha,\beta)=(0.4,1.0)$, supersonic propagation around $v_{\rm p}\approx 1.14$. At DNA-like values $(\alpha,\beta)\approx(1.2,1.1)$, the self-organized polarons remain bell-shaped and subsonic [1301.1851].

In field theory, anharmonic waves generalize the free plane wave by replacing $e^{iz}$ with
$$
\psi(z)=\exp[i(z+\varphi(z))],
$$
where $\varphi(z)$ is periodic and odd. The corresponding Fourier expansion contains harmonics of the phase variable and, in the most general class, may include a zero-frequency term interpreted as a non-zero vacuum expectation value. This usage again differs from narrow DSP notions of an inharmonic series model, but it preserves the core idea that exact interacting solutions need not remain purely sinusoidal [1108.1736].

Auditory modeling introduces yet another variant. A linear cochlear model made of damped strings keeps the mechanics linear but assumes that the quantity transmitted onward is stored energy. For a sinusoidal forcing at angular frequency $\Omega$, the $n$th string mode produces energy peaks near
$$
\xi=\frac{\Omega}{n},
$$
and, under a spatially uniform forcing assumption, only odd modes are driven. The resulting “sub-harmonic series” is therefore a pattern of energy maxima across place, not a physical subharmonic displacement spectrum. Combination-tone behavior then arises from the quadratic nature of the energy observable rather than from nonlinear mechanics in the displacement equation [2509.26395].

Taken together, the literature suggests that “inharmonic series model” is an umbrella term for several related but non-identical constructions. Their common feature is the relaxation of exact integer-multiple structure while retaining some nominal harmonic scaffold; what varies is whether the departure is encoded as deterministic frequency offsets, stochastic perturbations, waveform evolution, optimal-transport distance, modal detuning, or perturbative nonlinear correction.

Source: https://www.emergentmind.com/topics/inharmonic-series-model