Papers
Topics
Authors
Recent
Search
2000 character limit reached

Inharmonic Series Model in Signal Analysis

Updated 12 July 2026
  • Inharmonic series models represent signals as sums of sinusoids with frequencies near integer multiples of a nominal fundamental, with deviations quantified by Δk.
  • They encompass deterministic and stochastic perturbations, waveform-field formulations, and optimal transport techniques to approximate nearly harmonic structures.
  • These models enable improved pitch estimation, robust speech segregation, and analysis of nonlinear phenomena in both audio processing and physical systems.

An inharmonic series model is a representation of a signal or dynamical system in which prominent components are not constrained to lie exactly at integer multiples of a single fundamental. In signal analysis, the canonical almost-harmonic formulation is

xt=k=1Krkeiϕk+iωkt,ωk=kω0+Δk,x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad \omega_k=k\omega_0+\Delta_k,

where the Δk\Delta_k quantify departures from exact harmonicity. The literature does not treat this as a single universal construction. Instead, it includes deterministic perturbations of a harmonic lattice, definitions of a nominal fundamental based on waveform fit or spectral transport, stochastic perturbation models, and phase–waveform models that avoid explicit partial decomposition altogether (Elvander et al., 2020, 0911.5171).

1. Canonical almost-harmonic formulation

The basic signal-processing form of an inharmonic series model writes a finite-length or discrete-time signal as a sum of complex exponentials with frequencies near a harmonic scaffold,

xt=k=1Krkeiϕk+iωkt,t=0,1,,N1,x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad t=0,1,\ldots,N-1,

with

ωk=kω0+Δk.\omega_k=k\omega_0+\Delta_k.

In the perfectly harmonic case, all Δk\Delta_k vanish and ω0\omega_0 is an ordinary fundamental frequency. In the inharmonic case, the Δk\Delta_k are assumed small in the sense that Δkω0\Delta_k\ll\omega_0. This model is the standard “almost harmonic” representation: each partial is a perturbation of an integer multiple of a nominal base frequency rather than an unrelated sinusoid (Elvander et al., 2020).

The same framework admits structured special cases. A stiff-string example is

ωk=kω01+k2β,\omega_k=k\omega_0\sqrt{1+k^2\beta},

so that

Δk=kω0(1+k2β1)>0.\Delta_k=k\omega_0\left(\sqrt{1+k^2\beta}-1\right)>0.

Here the inharmonicity is systematic rather than arbitrary. A generic harmonic approximant used for analysis is

Δk\Delta_k0

with Δk\Delta_k1 not necessarily equal to Δk\Delta_k2. This approximant is central because many estimation procedures operate on the harmonic model even when the true signal is only almost harmonic (Elvander et al., 2020).

A plausible implication is that the phrase “inharmonic series model” is best understood as a family of nearby constructions centered on a nominal harmonic lattice, not as a single parametrization. Some models treat the deviations Δk\Delta_k3 as deterministic offsets, others as random variables, and others replace explicit partials by a different state space entirely.

2. Fundamental frequency in inharmonic signals

A central difficulty is that an inharmonic signal has no unique obvious notion of fundamental frequency. One formalization defines Δk\Delta_k4 as the best Δk\Delta_k5 harmonic approximation:

Δk\Delta_k6

Under a misspecified Gaussian harmonic model, this is exactly the pseudo-true parameter, namely the Kullback–Leibler best approximation inside the harmonic family. In that setting, the pseudo-true noise variance becomes

Δk\Delta_k7

This definition is useful as a benchmark for harmonic estimators applied to inharmonic data, but it depends on observation length and nuisance phases, and it can become ambiguous as Δk\Delta_k8 because different subharmonic candidates may fit equally well (Elvander et al., 2020).

A second definition replaces waveform fit by spectral transport. With random initial phases so that the process is wide-sense stationary, the inharmonic spectrum is

Δk\Delta_k9

and a harmonic spectrum is

xt=k=1Krkeiϕk+iωkt,t=0,1,,N1,x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad t=0,1,\ldots,N-1,0

The distance is defined by Monge–Kantorovich transport with cost

xt=k=1Krkeiϕk+iωkt,t=0,1,,N1,x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad t=0,1,\ldots,N-1,1

For fixed xt=k=1Krkeiϕk+iωkt,t=0,1,,N1,x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad t=0,1,\ldots,N-1,2, the optimization reduces to

xt=k=1Krkeiϕk+iωkt,t=0,1,,N1,x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad t=0,1,\ldots,N-1,3

For sufficiently small perturbations, the minimizing xt=k=1Krkeiϕk+iωkt,t=0,1,,N1,x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad t=0,1,\ldots,N-1,4 is explicitly

xt=k=1Krkeiϕk+iωkt,t=0,1,,N1,x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad t=0,1,\ldots,N-1,5

This is phase-free and observation-length-free, and it interprets the nominal fundamental as a power-weighted projection of perturbed partial frequencies onto the harmonic manifold (Elvander et al., 2020).

A third definition treats the inharmonic series as a randomly perturbed harmonic signal,

xt=k=1Krkeiϕk+iωkt,t=0,1,,N1,x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad t=0,1,\ldots,N-1,6

with independent zero-mean perturbations, often specialized to

xt=k=1Krkeiϕk+iωkt,t=0,1,,N1,x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad t=0,1,\ldots,N-1,7

Here xt=k=1Krkeiϕk+iωkt,t=0,1,,N1,x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad t=0,1,\ldots,N-1,8 is the mean harmonic scaffold, and the appropriate lower bound is a hybrid Cramér–Rao lower bound because the model mixes deterministic parameters with random latent perturbations. The associated ML/MAP criterion combines an unstructured sinusoidal data fit with a penalty on departures from a harmonic lattice. This yields a bridge between perfectly harmonic and fully unstructured sinusoidal models (Elvander et al., 2020).

3. Waveform-field and phase–time models

An alternative tradition does not represent a monophonic sound as a sum of partials at all. Instead, it models the signal as a function on a cylinder with independent shape-time and phase coordinates,

xt=k=1Krkeiϕk+iωkt,t=0,1,,N1,x_t=\sum_{k=1}^K r_k e^{i\phi_k+i\omega_k t},\qquad t=0,1,\ldots,N-1,9

where ωk=kω0+Δk.\omega_k=k\omega_0+\Delta_k.0. Ordinary audio is obtained by sampling the cylindrical field along a control path,

ωk=kω0+Δk.\omega_k=k\omega_0+\Delta_k.1

Here ωk=kω0+Δk.\omega_k=k\omega_0+\Delta_k.2 controls shape or timbral progression, while ωk=kω0+Δk.\omega_k=k\omega_0+\Delta_k.3 controls phase progression. The source signal is lifted to the cylinder by interpolation,

ωk=kω0+Δk.\omega_k=k\omega_0+\Delta_k.4

This construction is explicitly presented as robust against non-harmonic frequencies because it does not require decomposition into harmonically locked sinusoids (0911.5171).

For a sinusoidal input ωk=kω0+Δk.\omega_k=k\omega_0+\Delta_k.5, the model decomposes the frequency as

ωk=kω0+Δk.\omega_k=k\omega_0+\Delta_k.6

and, with Whittaker interpolation, the resynthesized frequency under shape-speed ωk=kω0+Δk.\omega_k=k\omega_0+\Delta_k.7 and phase-speed ωk=kω0+Δk.\omega_k=k\omega_0+\Delta_k.8 is

ωk=kω0+Δk.\omega_k=k\omega_0+\Delta_k.9

This preserves pure sines and gives a predictable mapping for non-integer frequencies without requiring a fixed harmonic series. In this representation, the integer part of a component is tied to phase progression and the residual detuning to shape-time progression. The model therefore encodes inharmonicity as residual detuning or waveform evolution across the shape axis rather than as an explicit bank of inharmonic partial tracks (0911.5171).

The cylindrical framework also proves exact structural properties. It is linear, it preserves a static periodic wave as a constant waveform around the cylinder, and it contains ordinary resampling as a special case. If the input has the separable form “envelope times waveform,” then the model preserves that factorization:

Δk\Delta_k0

hence

Δk\Delta_k1

Its stated limitations are equally clear: it is aimed at monophonic signals with a known constant period, it excludes noise portions such as speech in the exact theory, it is not polyphonic, and it does not preserve formant structure (0911.5171).

4. Estimation and inverse modeling

One important use of an inharmonic series model is not synthesis but estimation under model mismatch. A misspecified harmonic estimator assumes

Δk\Delta_k2

even when the true data come from an inharmonic process

Δk\Delta_k3

In the Gaussian setting, the pseudo-true parameter is the least-squares harmonic projection,

Δk\Delta_k4

and the pseudo-true variance is inflated by deterministic mismatch:

Δk\Delta_k5

The misspecified Gaussian Cramér–Rao lower bound then quantifies the bias–variance tradeoff of estimating an inharmonic signal through a harmonic scaffold. The stated conclusion is that, for moderately inharmonic signals, the misspecified harmonic model can achieve lower mean squared error than a correctly specified but unstructured sinusoidal model, and voiced speech is reported to belong to this class (Elvander et al., 2019).

A more recent inverse formulation keeps the signal approximately harmonic,

Δk\Delta_k6

but does not fit a parametric law for the Δk\Delta_k7. Instead, it represents spectral content as a measure on the unit circle and assigns spectral mass to candidate fundamentals through optimal transport. The normalized ground cost is

Δk\Delta_k8

Group sparsity over candidate fundamentals is enforced with

Δk\Delta_k9

This decouples the measurement operator from the harmonic prior: the forward model uses Fourier atoms, while harmonicity is enforced in the regularizer. The stated consequence is robustness to inharmonicity without introducing a stretched-comb law such as ω0\omega_00 (Björkman et al., 4 Aug 2025).

These estimation frameworks share a common geometry. The signal is not forced to be exactly harmonic; rather, it is projected, penalized, or softly clustered relative to a harmonic scaffold. This suggests that many practical “inharmonic series models” are best read as structured approximations, not as claims of exact periodicity.

5. Audio, speech, and perceptual consequences

The computational importance of harmonic versus inharmonic structure is especially visible in monaural speech segregation. One study defines harmonic speech as having Fourier components at the fundamental frequency ω0\omega_01 and its integer multiples, and generates inharmonic speech by jittering each harmonic:

ω0\omega_02

The tested range is

ω0\omega_03

Under this slight departure from exact harmonicity, end-to-end separators degrade sharply: Conv-TasNet is reported to drop from ω0\omega_04 dB to ω0\omega_05 dB under ω0\omega_06 harmonic jitter, and DPT-Net from ω0\omega_07 dB to ω0\omega_08 dB for imperceptible amounts of harmonic jitter. Training on inharmonic speech does not remove the sensitivity and worsens natural-speech performance. The paper interprets this as evidence that these networks rely strongly on harmonic spectral patterns, whereas temporal-coherence systems rely primarily on timing cues and are much less tied to exact harmonic spacing (Parikh et al., 2022).

Inharmonic series models also matter for roughness and consonance calculations. A Helmholtz/Plomp–Levelt dissonance formulation operates directly on arbitrary partial frequencies and amplitudes. For a pair of partials,

ω0\omega_09

and the dyadic roughness is

Δk\Delta_k0

with Δk\Delta_k1 for Δk\Delta_k2. Total dissonance is not obtained by naive summation but by a loudness/intensity-weighted aggregate, with preferred exponent Δk\Delta_k3. Because the formulas are expressed in terms of arbitrary partial frequencies Δk\Delta_k4 and intensities, the method can be applied directly to inharmonic spectra rather than only to exact harmonic tones (Dillon, 2013).

Together these results indicate two distinct uses of the concept. In one use, inharmonicity is a nuisance that destabilizes systems trained on exact harmonic structure. In the other, it is the primary object of analysis, as in roughness computation or robust pitch estimation.

6. Broader physical and mathematical extensions

Outside signal processing, related models appear in nonlinear dynamics, mathematical physics, and auditory modeling. In a delay-differential toy model of flute-like instruments, the resonator is written as

Δk\Delta_k5

and inharmonicity is controlled by the ratio Δk\Delta_k6. The cases Δk\Delta_k7, Δk\Delta_k8, and Δk\Delta_k9 produce qualitatively different behavior: bistable hysteresis, an intermediate quasiperiodic regime, or first-register dominance. Here an “inharmonic series model” means a low-order modal resonance series whose detuning actively changes nonlinear regime selection (Terrien et al., 2012).

For a periodically forced anharmonic chain, the asymptotic periodic state is represented as a convergent power series in the anharmonicity parameter,

Δkω0\Delta_k\ll\omega_00

The expansion is proved to converge for Δkω0\Delta_k\ll\omega_01 under the nonresonance condition that no integer multiple of the forcing frequency falls into the harmonic spectral band

Δkω0\Delta_k\ll\omega_02

In this setting, “anharmonic series model” refers not to detuned partials but to a perturbation series for the unique periodic solution of a nonlinear lattice (Garrido et al., 30 Mar 2025).

A one-dimensional polaron model on an inharmonic lattice uses the cubic FPU-type potential

Δkω0\Delta_k\ll\omega_03

Its continuum equations admit exact one-soliton solutions only on a special relation between the electron–phonon and inharmonicity parameters, while numerical simulations in a stronger inharmonic regime produce unusual stable moving polarons with multi-peak envelopes and, for Δkω0\Delta_k\ll\omega_04, supersonic propagation around Δkω0\Delta_k\ll\omega_05. At DNA-like values Δkω0\Delta_k\ll\omega_06, the self-organized polarons remain bell-shaped and subsonic (Astakhova et al., 2013).

In field theory, anharmonic waves generalize the free plane wave by replacing Δkω0\Delta_k\ll\omega_07 with

Δkω0\Delta_k\ll\omega_08

where Δkω0\Delta_k\ll\omega_09 is periodic and odd. The corresponding Fourier expansion contains harmonics of the phase variable and, in the most general class, may include a zero-frequency term interpreted as a non-zero vacuum expectation value. This usage again differs from narrow DSP notions of an inharmonic series model, but it preserves the core idea that exact interacting solutions need not remain purely sinusoidal (Himpsel, 2011).

Auditory modeling introduces yet another variant. A linear cochlear model made of damped strings keeps the mechanics linear but assumes that the quantity transmitted onward is stored energy. For a sinusoidal forcing at angular frequency ωk=kω01+k2β,\omega_k=k\omega_0\sqrt{1+k^2\beta},0, the ωk=kω01+k2β,\omega_k=k\omega_0\sqrt{1+k^2\beta},1th string mode produces energy peaks near

ωk=kω01+k2β,\omega_k=k\omega_0\sqrt{1+k^2\beta},2

and, under a spatially uniform forcing assumption, only odd modes are driven. The resulting “sub-harmonic series” is therefore a pattern of energy maxima across place, not a physical subharmonic displacement spectrum. Combination-tone behavior then arises from the quadratic nature of the energy observable rather than from nonlinear mechanics in the displacement equation (Boscain et al., 30 Sep 2025).

Taken together, the literature suggests that “inharmonic series model” is an umbrella term for several related but non-identical constructions. Their common feature is the relaxation of exact integer-multiple structure while retaining some nominal harmonic scaffold; what varies is whether the departure is encoded as deterministic frequency offsets, stochastic perturbations, waveform evolution, optimal-transport distance, modal detuning, or perturbative nonlinear correction.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Inharmonic Series Model.