Papers
Topics
Authors
Recent
Search
2000 character limit reached

ARTI-6: Six-Dimensional Speech Encoding

Updated 12 July 2026
  • ARTI-6 is a six-dimensional articulatory speech encoding framework that represents speech production through six specific vocal-tract regions, including the velum, tongue root, and larynx.
  • It derives continuous articulatory measures from real-time MRI, using fixed ROI analysis to capture physiologically grounded constriction actions.
  • The framework couples efficient acoustic-to-articulatory inversion with HiFi-GAN-based synthesis, achieving high intelligibility and computational efficiency for applications like low-bitrate codecs.

Searching arXiv for the ARTI-6 paper and a few directly related works mentioned in the provided data so the article can be grounded in current arXiv records. ARTI-6 is a compact six-dimensional articulatory speech encoding framework derived from real-time MRI data that represents speech production through six vocal-tract regions and couples that representation to both articulatory inversion and articulatory synthesis. Its design centers on three components: a six-dimensional articulatory feature set, an inversion model that predicts articulatory features from speech acoustics by leveraging speech foundation models, and an overview model that reconstructs intelligible speech directly from articulatory features. The framework is presented as interpretable, computationally efficient, and physiologically grounded, while explicitly extending articulatory coverage to the velum, tongue root, and larynx (Lee et al., 25 Sep 2025).

1. Conceptual scope and representational design

ARTI-6 encodes speech articulation with six features, each corresponding to a critical region of the vocal tract. The six regions of interest are Lip Aperture (LA), Tongue Tip (TT), Tongue Body (TB), Velum (VL), Tongue Root (TR), and Larynx (LX). The feature set is described as knowledge-driven and intended to cover essential constriction actions necessary for speech, with full vocal tract coverage unlike previous EMA-based approaches (Lee et al., 25 Sep 2025).

The choice of regions is motivated by articulatory phonology and speech science theory. In the formulation reported for ARTI-6, these regions capture both oral and nasal gestures and account for the main constriction sites relevant to speech sounds. A central distinction from earlier EMA-based feature sets is that ARTI-6 includes the velum, larynx, and tongue root, which the paper identifies as crucial for fuller coverage, especially for nasality and voicing (Lee et al., 25 Sep 2025).

Feature Region Functional description
LA Lip Aperture Degree of mouth opening
TT Tongue Tip Position and constriction degree
TB Tongue Body Position and constriction
VL Velum State of the soft palate
TR Tongue Root Constriction near the tongue root
LX Larynx State or position relevant to voicing

The framework therefore treats articulatory encoding not as an opaque latent space but as a small set of physiologically specified variables. Each dimension is tied to a concrete speech organ or constriction location rather than to an abstract acoustic embedding (Lee et al., 25 Sep 2025).

2. Derivation from real-time MRI

The articulatory representation is derived from real-time MRI videos of speech production. For each region of interest, ARTI-6 computes constriction degree by measuring the mean pixel intensity at a fixed ROI location in each rtMRI frame, thereby producing a continuous articulatory measure (Lee et al., 25 Sep 2025).

This extraction strategy gives the framework its physiological grounding. Because the representation is built from rtMRI rather than inferred indirectly from sparse sensor trajectories, the paper positions ARTI-6 as covering vocal-tract regions that are difficult to access in conventional EMA-style configurations, especially the velum, tongue root, and larynx (Lee et al., 25 Sep 2025).

A plausible implication is that the framework prioritizes anatomically motivated observables over purely data-driven compression. The paper’s emphasis on fixed ROI locations and continuous frame-level measures is consistent with a representation designed for interpretable articulatory trajectories rather than for unconstrained latent coding.

3. Acoustic-to-articulatory inversion

The articulatory inversion component predicts the six-dimensional articulatory vector directly from speech acoustics. The reported backbones are speech foundation models including WavLM, HuBERT, Whisper-Large, and Wav2Vec2-XLSR. Fine-tuning is performed with LoRA inserted into Transformer layers. The architecture combines outputs from convolutional and Transformer blocks through a linear layer, followed by a 1D-pointwise convolution and projection that yield six predicted articulatory feature values per frame (Lee et al., 25 Sep 2025).

The inversion objective is an ℓ2\ell_2 regression loss,

Linv=∥ypred−ytrue∥22.\mathcal{L}_{\text{inv}} = \| \mathbf{y}_{\text{pred}} - \mathbf{y}_{\text{true}} \|_2^2 .

Evaluation is reported using frame-wise Pearson correlation between predicted and ground-truth articulatory features. The best test-set prediction correlation is 0.872 with WavLM or HuBERT as backbone; Whisper-Large reaches 0.844, and Wav2Vec2-XLSR reaches 0.615. The abstract summarizes the inversion performance as a prediction correlation of 0.87. The paper also notes that accuracy is high for lips and tongue features and lower for velum and larynx, attributing this to imaging and noise issues (Lee et al., 25 Sep 2025).

These results position the inversion model as a compact acoustic-to-articulatory mapper rather than a full biomechanical estimator. The reported asymmetry across articulators suggests that the six-dimensional code is not uniformly easy to recover from acoustics, particularly for articulators whose rtMRI signatures are noisier or less directly reflected in the acoustic signal.

4. Articulatory-to-acoustic synthesis

The synthesis component reconstructs speech directly from the six articulatory features. The reported synthesizer is based on HiFi-GAN v1 and is conditioned on speaker embeddings extracted with ECAPA-TDNN to support multi-speaker synthesis (Lee et al., 25 Sep 2025).

The training objective is given as

Lgen=λ1LGAN+λ2Lrecon+λ3Lfeature-match.\mathcal{L}_{\text{gen}} = \lambda_1 \mathcal{L}_{\text{GAN}} + \lambda_2 \mathcal{L}_{\text{recon}} + \lambda_3 \mathcal{L}_{\text{feature-match}} .

The objective quality metrics reported for synthesis are a Word Error Rate of 12.5%, a Character Error Rate of 7.4%, and a UTMOS predicted MOS of 3.84. The subjective human MOS is 3.95. The paper states that these results show that even a low-dimensional representation can generate natural-sounding speech and that intelligibility and naturalness remain high despite the six-dimensional bottleneck (Lee et al., 25 Sep 2025).

The comparison reported in the paper is explicit: mel-spectrogram and EMA-based systems reach MOS values of approximately 4.4 and 4.3, with lower WER and CER, whereas ARTI-6 remains highly intelligible and fairly natural despite being only 6D. This suggests a trade-off between extreme compactness and maximal synthesis fidelity, rather than a claim that six dimensions saturate the speech production space.

5. Relation to other speech feature spaces

ARTI-6 is contrasted with several established feature spaces. The paper lists MFCC at 10–20 dimensions, mel-spectrogram at 80 dimensions, latent representations such as WavLM at 200+ dimensions, and EMA plus pitch plus loudness at 14 dimensions. Against these alternatives, ARTI-6 is characterized as six-dimensional, interpretable, knowledge-based, and directly covering the velum, tongue root, and larynx (Lee et al., 25 Sep 2025).

This comparison sharpens the framework’s intended niche. MFCC, mel-spectrogram, and high-dimensional latent spaces are perceptual or learned acoustic encodings rather than directly articulatory descriptions. EMA-based representations are articulatory, but the paper argues that they do not directly cover the velum, tongue root, and larynx. ARTI-6 therefore presents itself as a compact articulatory code with broader physiological coverage than prior low-dimensional articulatory alternatives (Lee et al., 25 Sep 2025).

A common misconception would be to treat ARTI-6 as simply another low-bitrate acoustic representation. The paper instead frames it as an articulatory encoding whose interpretability is intrinsic: each dimension corresponds to a specified vocal-tract region. Its compactness is therefore coupled to anatomical semantics, not merely to compression.

6. Significance, applications, and availability

The significance assigned to ARTI-6 in the paper rests on three properties. First, interpretability: each dimension corresponds to a concrete speech organ or articulatory location. Second, compactness: the 6D vector is contrasted with representations having dozens or hundreds of dimensions, which the paper associates with reduced computation, reduced memory, and low-latency speech transmission and processing. Third, physiological fidelity: the representation is positioned as useful for modeling and analyzing speech production when imaging data is unavailable (Lee et al., 25 Sep 2025).

The application space identified for the framework includes articulatory inversion, articulatory synthesis, silent speech interfaces, assistive technologies, low-bitrate speech codecs, and brain–speech or muscle–speech studies. The paper therefore places ARTI-6 at the intersection of articulatory phonetics, speech representation learning, and generative speech technology (Lee et al., 25 Sep 2025).

The source code and speech samples are publicly available. This availability is consequential because the framework combines a physiologically grounded feature definition, a foundation-model-based inversion model, and a HiFi-GAN-based synthesis model in a single encoding pipeline. A plausible implication is that ARTI-6 can function as an intermediate articulatory code for systems that require both interpretability and computational economy, while accepting a measurable but bounded trade-off relative to higher-dimensional acoustic representations.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ARTI-6.