Papers
Topics
Authors
Recent
Search
2000 character limit reached

Explicit Articulatory Control

Updated 14 July 2026
  • Explicit Articulatory Control is defined by interpretable articulatory parameters (e.g., vocal-tract geometry, gestural targets) that directly map to the physical speech apparatus.
  • It integrates diverse methodologies such as nonlinear task-dynamic equations, gesture-score models, differentiable physical synthesis, and neural articulatory codes for enhanced speech modeling.
  • The approach facilitates practical applications including human and reinforcement-learning interfaces, improved speech inversion, and robust speaker-specific articulatory modeling.

Searching arXiv for papers on explicit articulatory control and closely related articulatory modeling, inversion, and synthesis. Explicit articulatory control denotes a family of speech and singing control formulations in which the primary variables are articulatory, gestural, kinematic, or physically interpretable parameters—such as task variables, vocal-tract geometry, EMA trajectories, or region-specific articulatory codes—rather than opaque latent units or directly specified spectral envelopes. Across recent work, the term spans nonlinear task-dynamic equations, gesture-score models, differentiable physical synthesis, articulatory inversion and coding, and human or reinforcement-learning interfaces. What unifies these approaches is that the control parameters retain a clear relation to the vocal apparatus and its acoustics (Cámara et al., 3 Jun 2026, Kirkham, 2024, Lee et al., 25 Sep 2025, Cho et al., 2024).

1. Conceptual scope and defining properties

In current research, explicit articulatory control is not a single architecture but a design principle. In the sygyt copy-synthesis study, “explicit” means that the primary control variables are vocal-tract geometry and physical parameters—tube areas, a sublingual side-branch, spatially varying damping, and LF source parameters—optimized end-to-end by gradient descent while the structure remains a physical Kelly–Lochbaum waveguide; this is contrasted directly with DDSP-style per-harmonic amplitude control that has no vocal-tract interpretation (Cámara et al., 3 Jun 2026). In nonlinear task-dynamic work, explicitness instead resides in the control law itself: targets, stiffness, damping, and nonlinear restoring-force parameters are written as explicit equations in task space, with scaling laws added so that parameters remain interpretable across movement distances (Kirkham, 2024).

This suggests that explicitness is best understood as semantic transparency of the control variables rather than commitment to one representational level. In some models, the explicit variables are task variables such as constriction degree or lip aperture. In others, they are gestural targets and activation intervals, articulatory trajectories inferred from acoustics, or directly manipulable physical geometry. The same principle also appears in low-dimensional articulatory coding: ARTI-6 defines six scalar dimensions, each corresponding to a specific vocal-tract subsystem—lip aperture, tongue tip, tongue body, velum, tongue root, and larynx—so that each dimension functions as a continuous control knob (Lee et al., 25 Sep 2025). SPARC similarly treats vocal-tract kinematics and source features as a 14-dimensional control stream, with a separate 64-dimensional speaker embedding for voice texture (Cho et al., 2024).

A useful way to organize the literature is by the level at which control variables are made explicit.

Control level Representative variables Representative papers
Task dynamics x,T,k,m,b,dx, T, k, m, b, d (Kirkham, 2024)
Gesture and tract variables vocalic height, front–back, lip rounding, constriction degree, velum and glottis state (Kröger, 11 Nov 2025, Kröger, 27 Jul 2025)
Kinematic codes LA, TT, TB, VL, TR, LX; EMA trajectories; articulatory + source features (Lee et al., 25 Sep 2025, Cho et al., 2024, Kim et al., 15 Jun 2026)
Physical synthesis parameters tube areas, damping, LF source parameters, side-branches (Cámara et al., 3 Jun 2026)

2. Task-space, gesture-space, and tract-variable formulations

Within task-dynamic theories of speech, control is explicit in task space rather than in muscle coordinates. The canonical equation is the critically damped mass–spring system

mx¨+bx˙+k(x−T)=0,m\ddot{x} + b\dot{x} + k(x - T) = 0,

where xx is the task variable, TT the target, mm inertia, bb damping, and kk stiffness (Kirkham, 2024). To address empirical shortcomings of linear dynamics, the nonlinear extension

mx¨+bx˙+kx−dx3=0m\ddot{x} + b\dot{x} + kx - dx^3 = 0

adds a cubic restoring force. The paper’s central contribution is that the nonlinear coefficient must be scaled if it is to remain interpretable across movement amplitudes. The basic inverse-square scaling is

d′=dk∣x0−T∣2,d' = \frac{dk}{|x_0 - T|^2},

and the generalized version introduces a global movement range DD and

mx¨+bx˙+k(x−T)=0,m\ddot{x} + b\dot{x} + k(x - T) = 0,0

In this formulation, mx¨+bx˙+k(x−T)=0,m\ddot{x} + b\dot{x} + k(x - T) = 0,1 governs timing and peak speed, while mx¨+bx˙+k(x−T)=0,m\ddot{x} + b\dot{x} + k(x - T) = 0,2 becomes a bounded measure of kinematic nonlinearity and velocity-profile symmetry (Kirkham, 2024).

A different but related explicitness appears in DYNARTmo, where speech is represented as a gesture score distributed across vocalic, consonantal, velopharyngeal, glottal, and pulmonary tiers. Each gesture specifies onset and offset times, rise and fall durations, a target vector, and a pull value. The gesture activation is cosine-shaped, and overlapping gestures affecting the same control parameter are blended by

mx¨+bx˙+k(x−T)=0,m\ddot{x} + b\dot{x} + k(x - T) = 0,3

The resulting control variables remain interpretable as vocalic height, fronting, lip rounding, constriction degree, velum position, or glottal aperture (Kröger, 11 Nov 2025). The static DYNARTmo formulation complements this with ten continuous parameters and six discrete parameters controlling six articulators—lips, tongue tip, tongue dorsum, lower jaw, velum, and glottis—via vowel-space interpolation, consonantal constriction parameters, and derived jaw adjustments (Kröger, 27 Jul 2025).

These formulations differ in mechanics but share the same structural property: they turn speech control into the evolution of a small set of named variables with explicit dynamics and explicit mappings to articulator geometry.

3. Differentiable physical synthesis and geometry-level control

The most direct realization of explicit articulatory control is to let the control variables be the parameters of a physical vocal-tract model. The sygyt study on biphonic singing exemplifies this strategy. It models the vocal system as a Kelly–Lochbaum digital waveguide with 44 oral segments, 28 nasal segments, and a 15-segment sublingual tube, running at 16 kHz, with two three-way junctions at the sublingual junction and velum (Cámara et al., 3 Jun 2026). Each segment mx¨+bx˙+k(x−T)=0,m\ddot{x} + b\dot{x} + k(x - T) = 0,4 has a cross-sectional area mx¨+bx˙+k(x−T)=0,m\ddot{x} + b\dot{x} + k(x - T) = 0,5 and damping coefficient mx¨+bx˙+k(x−T)=0,m\ddot{x} + b\dot{x} + k(x - T) = 0,6, and the reflection coefficient at a boundary is

mx¨+bx˙+k(x−T)=0,m\ddot{x} + b\dot{x} + k(x - T) = 0,7

After scattering, spatially varying damping is applied as

mx¨+bx˙+k(x−T)=0,m\ddot{x} + b\dot{x} + k(x - T) = 0,8

with mx¨+bx˙+k(x−T)=0,m\ddot{x} + b\dot{x} + k(x - T) = 0,9. Values near xx0 produce high-xx1 resonances; values nearer xx2 broaden them.

The central control mechanism is a cubic B-spline parameterization of oral and sublingual diameter and damping profiles. With control points xx3 and xx4,

xx5

For the oral tract, xx6 and xx7. This produces about 70 DOFs/frame for the B-spline model, compared with about 13 DOFs/frame for the articulator-chain baseline and about 318 parameters per frame for the DDSP harmonic-plus-noise baseline (Cámara et al., 3 Jun 2026).

The model also adds a sublingual secondary LF source injected at oral segment 9, driven at the overtone frequency xx8, with independent amplitude, tensiveness, open quotient offset, and spectral tilt offset. In ablation, removing the sublingual source increases LSD from 9.34 to 10.32 dB across 20 segments, whereas removing spatially varying damping increases LSD by only xx9 dB, indicating that the secondary source is the dominant contributor among the added articulatory features under the reported metrics (Cámara et al., 3 Jun 2026).

Optimization is fully differentiable in JAX and uses a compound objective

TT0

with all TT1 terms equal to 1. The overtone-salience term

TT2

directly targets sygyt’s salient overtone. On 20 segments from two datasets, the proposed model reduces log-spectral distance by 30–38% relative to an articulatory baseline—13.84 dB to 9.64 dB on HFA and 14.53 dB to 9.04 dB on Bergevin—and yields mean formant peak error of 28 Hz in the 1–3 kHz overtone region, versus 222 Hz for the articulator chain and 120 Hz for DDSP (Cámara et al., 3 Jun 2026).

This case is important because it shows explicit articulatory control at the geometry-and-physics level, not only in kinematic or symbolic form. The controllable parameters are directly interpretable as cavity shape, coupling, damping, and source settings, yet remain optimizable from audio.

4. Learned articulatory codes, inversion, and articulatory coding systems

A large branch of the literature seeks explicit articulatory control by learning compact, interpretable codes from imaging or acoustics. ARTI-6 is a particularly clear example. It defines a six-dimensional articulatory encoding from real-time MRI using fixed ROIs for LA, TT, TB, VL, TR, and LX, with each feature computed as average pixel intensity inside the corresponding ROI. The articulatory inversion model, built by fine-tuning speech foundation models with LoRA, achieves an average correlation of 0.872 on the test set with WavLM-Large; the articulatory synthesis model, based on HiFi-GAN v1 and conditioned on ARTI-6 trajectories plus ECAPA-TDNN speaker embeddings, yields WER 0.125, CER 0.074, UTMOS 3.840, and MOS 3.947 on LibriTTS-R test-clean (Lee et al., 25 Sep 2025). The framework is explicitly designed so that each dimension is interpretable and directly manipulable.

MirrorNet approaches explicit control from a sensorimotor-learning perspective. Its latent representation is explicitly sized to 9 articulatory parameters—six vocal tract variables plus aperiodicity, periodicity, and pitch—and it is coupled to a frozen articulatory synthesizer. A brief initialization phase uses about 30 minutes of articulatory data to anchor the latent space to true TVs and source parameters; subsequent learning is unsupervised. With initialization and the fully trained synthesizer, MirrorNet reaches average correlation 0.8286 over six TVs and 0.8540 over all nine parameters, compared with 0.7848 for a fully supervised BiGRNN baseline. Without initialization, performance collapses to 0.4420 over the six TVs, showing that articulatory explicitness does not emerge reliably unless the latent space is anchored to measured control variables (Siriwardena et al., 2022).

SPARC generalizes the same idea into a neural articulatory codec. Its frame-level control stream is 14-dimensional—12 EMA-style articulator coordinates plus pitch and loudness—at 50 Hz, with a separate 64-dimensional speaker embedding. The articulatory analysis model uses WavLM Large; the articulatory synthesis model is a HiFi-GAN-style vocoder with FiLM conditioning by speaker identity. On LibriTTS-R test-clean, ground truth WER 4.22% becomes 5.43% after resynthesis, with MOS 3.82 for resynthesized speech; in zero-shot voice conversion, articulation coding–recoding PCC is TT3, pitch PCC 0.944, loudness PCC 0.938, speaker similarity 0.879, and SID accuracy 94.6% among 107 VCTK speakers (Cho et al., 2024). This makes the articulatory code simultaneously interpretable, usable for explicit editing, and separable from speaker identity.

Data scarcity remains a central obstacle. ArtBoost addresses it by extracting pseudo articulatory trajectories for visible articulators from the TFHP speech–mesh dataset and pre-training AAI models before fine-tuning on real EMA. On HPRC, SSL-AAI improves from PCC 0.678 to 0.698 and RMSE 0.736 to 0.717; on USC-TIMIT, PCC improves from 0.351 to 0.510 and RMSE from 0.864 to 0.792. The gains transfer across both SSL-AAI and SI-AAI, suggesting that explicit articulatory trajectories can be made more reliable by large-scale pseudo-supervision even when only visible articulators are directly observed (Kim et al., 15 Jun 2026).

A related question is whether articulatory spaces remain consistent across speakers. Using WavLM-based AAI, minimal-pair target extraction, and cross-speaker classification, work on interspeaker consistency reports that real XRMB articulatory targets reach 92% classification accuracy, while learned spaces can be improved further by speech-only consistency training on paired utterances with identical text. In English, this consistency training improves minimal-pair classification for the LoRA model and also improves WER in articulatory synthesis; however, Russian performance drops, indicating that cross-speaker consistency and cross-linguistic consistency are not identical objectives (McGhee et al., 26 May 2025).

5. Human interfaces, non-speech controllers, and policy learning

Explicit articulatory control also appears in systems where a human or an RL policy directly drives articulatory variables. In “Sound-Stream II,” four continuous force-controlled DOFs are mapped to excitations of superior longitudinal, inferior longitudinal, anterior genioglossus, and posterior genioglossus in a 2D ArtiSynth tongue. The chain is explicit end to end: TT4 A 22-section area function is computed from tongue–palate distances and sent in real time to JASS, with Rosenberg-type glottal excitation (Saha et al., 2018).

“SPEAK WITH YOUR HANDS” replaces muscle-level control with a hand-kinematic interface. CyberGlove II provides 18 sensor signals; wrist flexion and deviation control global tongue height and protrusion, while finger motions shape local constrictions. These sensor values define a spline tongue model, the difference between spline tongue and fixed palate is used as diameter, and the resulting 1D area function drives Pink Trombone in real time (Saha et al., 2021). The control space remains articulatory: the user does not specify phonemes or formants, but tongue geometry.

A more autonomous control regime appears in reinforcement-learning work on articulatory speech generation. There, speech production is cast as a motor-control problem in which a PPO policy outputs 13 continuous controls: 2D velocities for tongue dorsum, tongue blade, tongue tip, lower incisor, upper lip, and lower lip, plus loudness. Each episode lasts 50 timesteps, approximately one second of speech. SPARC decodes articulatory trajectories into audio, and Sylber provides syllable-level acoustic feedback; the reward is cosine similarity between the policy-generated syllable embedding and the target embedding, with a penalty of TT5 if no syllable is detected. Trained separately on six syllables, the framework achieves similarity scores above 0.85 for several syllables and correct human transcription for “please,” “loot,” and “cat” (Anand et al., 7 Oct 2025). This is explicit articulatory control in the strong sense that every action is a directly interpretable articulator velocity.

Transformer-based phoneme-to-articulatory estimation provides a complementary route to explicit temporal control. In the FastSpeech-based PTA model, a duration predictor and length regulator create hard alignments between phoneme sequences and articulatory frames, turning durations into explicit control variables. Compared with Tacotron, the model yields relative CC improvements of 154% for subject-dependent, 11.8% for pooled, and 4.8% for fine-tuning strategies (Udupa et al., 2021). A plausible implication is that explicit articulatory control benefits not only from explicit spatial variables but also from explicit timing variables.

6. Interpretability, sources of variation, and limits of the paradigm

A recurrent claim across the literature is that explicit articulatory control provides interpretability that acoustic-latent systems do not. In the sygyt waveguide, learned diameter and damping profiles show constrictions near the sublingual junction in 16/20 segments and a pattern of low damping near the junction but high damping in the anterior oral cavity, consistent with overtone-focusing descriptions from MRI and acoustic studies (Cámara et al., 3 Jun 2026). In DYNARTmo, control parameters map directly onto visible sagittal, glottal, and palatal views, which makes the model suitable for phonetics education and speech therapy (Kröger, 27 Jul 2025). In ARTI-6, each scalar dimension corresponds to a named subsystem of the vocal tract, while in SPARC the articulatory code can be rendered as animated kinematics (Lee et al., 25 Sep 2025, Cho et al., 2024).

At the same time, the literature also shows that articulatory explicitness does not eliminate variability. A study of Northern-Anglo English /i/ demonstrates that speaker-specific tongue-shape strategies for palatal vowels systematically predict formant dynamics in /i, eɪ, aɪ, ɔɪ/. Using ultrasound tongue imaging from 36 speakers, the study identifies three principal shape components for /i/, then shows via GAMMs that these speaker-level strategy variables predict F1 and F2 trajectory timing and slope. Greater articulatory displacement of tongue root and/or dorsum produces greater distortion from the mean tongue shape in palatal vowels and also requires higher articulatory velocities, resulting in relatively earlier and steeper formant transitions (Strycharczuk et al., 22 May 2026). This suggests that explicit articulatory control is not merely a control interface; it is also a framework for explaining structured speaker variation.

The same distinction is important at the boundary between articulatory and non-articulatory explicitness. GTR-Voice introduces an explicit control space for glottalization, tenseness, and resonance, with 125 legal GTR combinations recorded over 20 Chinese sentences. FastPitch and StyleTTS conditioned on these labels allow listeners to identify intended changes along each dimension above chance, with overall controllability accuracies of 67.07% and 57.14% respectively (Li et al., 2024). WordVoice similarly introduces explicit duration, boundary, energy, pitch, and tone control at the word level, but it is careful to frame this as explicit prosodic/acoustic planning rather than direct articulator control (Nie et al., 7 Jul 2026). These systems are adjacent to explicit articulatory control but not equivalent to it.

Current limitations are consistent across papers. The sygyt work is restricted to one singing style, fixes TT6 and overtone tracks during optimization, and requires about 30 minutes CPU per 5-second segment (Cámara et al., 3 Jun 2026). ARTI-6 derives its six-dimensional space from a single rtMRI speaker and reports lower reliability for velum and larynx features (Lee et al., 25 Sep 2025). ArtBoost provides pseudo supervision only for visible articulators, not internal tongue sensors (Kim et al., 15 Jun 2026). DYNARTmo’s current static model does not yet publish the full dynamic movement generation and articulatory–acoustic coupling that it is designed to support (Kröger, 27 Jul 2025). Reinforcement-learned articulatory control currently uses separate syllable-specific policies rather than a single goal-conditioned speech policy (Anand et al., 7 Oct 2025).

Taken together, these results support a broad but precise interpretation of explicit articulatory control. It is not confined to one representational regime, but it always requires that the operative variables be physically or functionally interpretable and that they stand in an explicit relation to the speech apparatus. Physical waveguides, task-dynamic equations, gesture scores, MRI- or EMA-derived codes, and direct human or policy interfaces all satisfy this requirement in different ways. A plausible implication is that future progress will come less from replacing one explicit formalism with another than from integrating them: stable articulatory codes, explicit gesture timing, differentiable physical synthesis, and speaker-robust training within a common control architecture.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Explicit Articulatory Control.