---
title: Conformer-based Ultrasound-to-Speech Conversion
url: https://www.emergentmind.com/topics/conformer-based-ultrasound-to-speech-conversion
type: topic
---

# Conformer-based Ultrasound-to-Speech Conversion

ChordFormer is a conformer-based neural architecture for large-vocabulary audio chord recognition, specifically designed to address the challenges arising from the long-tail distribution of chord types, the need to model both local spectral detail and long-range harmonic context, and the requirement for structured, musically meaningful chord representations. ChordFormer combines a constant-Q spectrogram frontend, a stack of conformer blocks (hybridizing multi-head self-attention with convolutional modules), and a chain-conditional random field (CRF) decoder. The system outputs a joint prediction over six structured chord components, facilitating precise recognition of complex chord types even in the presence of significant class imbalance [2502.11840].

## 1. Motivation and Design Objectives

Chord recognition in music information retrieval involves mapping audio to symbolic chord labels that reflect musical structure. Existing systems have achieved robust accuracy for simple chord vocabularies (e.g., major/minor triads), but large-vocabulary recognition—covering rare extensions and inversions—remains challenging. The rare occurrence of many chord types in available datasets results in pronounced class imbalance, hindering effective learning. ChordFormer is designed to:

- Transcribe audio to rich, structurally decomposed chord labels (root+triad, bass, extensions up to 13th).
- Mitigate the long-tail distribution by leveraging a reweighted loss for underrepresented classes.
- Model both local audio features (voicing, partials) and global dependencies (harmonic progressions).
- Provide structured outputs amenable to music-theoretic interpretation and joint learning across chord components [2502.11840].

## 2. Frontend and Feature Representation

ChordFormer accepts audio signals sampled at 22 050 Hz. Inputs are transformed using the Constant-Q Transform (CQT), spanning musical range C1–C8 with 36 bins per octave, resulting in 252 frequency bins per frame. The time step is set by a hop length of 512 samples (approximately 23.2 ms). Spectrogram values are converted to decibel scale using amplitude-to-db normalization (Librosa), then globally normalized. Data augmentation employs pitch shifting by –5 to +6 semitones, applied consistently to both the input spectrogram and the associated chord labels.

For each frame $t$, the chord is represented by a 6-dimensional vector:
$$
Z^{(t)} = [z_1^{(t)}, z_2^{(t)}, ..., z_6^{(t)}]
$$
where each $z_j^{(t)}$ encodes a musically meaningful component:

| Component ($z_j$)      | Category      | Classes                           |
|------------------------|--------------|-----------------------------------|
| $z_1$: root+triad      | 13 roots × 7 triads + "N" | 92 (including "no-chord")   |
| $z_2$: bass            | 12 chroma + "N"           | 13                            |
| $z_3$: 7th extension   | N, 7, ♭7, ♭♭7             | 4                             |
| $z_4$: 9th extension   | N, 9, ♯9, ♭9              | 4                             |
| $z_5$: 11th extension  | N, 11, ♯11                | 3                             |
| $z_6$: 13th extension  | N, 13, ♭13                | 3                             |

This structured target reduces the large-vocabulary classification problem to six smaller multiclass subproblems, reflecting hierarchical relationships in music theory and facilitating parameter sharing [2502.11840].

## 3. Conformer Block Architecture

The core of ChordFormer consists of $N=4$ conformer blocks operating on a feature dimension $D=256$. The conformer block fuses local and global modeling via the following sequence:

1. **Input Projection:** The CQT spectrogram (252 bins) is linearly projected to 256 dimensions.
   
2. **Half-Step Feed-Forward Module:** A position-wise feed-forward network (FFN) with Swish activation, pre-norm, dropout, and a half-step residual connection:
   $$
   \tilde Z_i = Z_i + \frac{1}{2} \cdot \text{FFN}(Z_i)
   $$
3. **Multi-Head Self-Attention (MHSA):** With relative sinusoidal positional encodings. For $n_h=4$ heads:
   $$
   \text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_K}}\right)V
   $$
   The output is added residually:
   $$
   Z_i^{(a)} = \tilde Z_i + \text{MHSA}(\tilde Z_i)
   $$
4. **Convolution Module:** Pre-normed, pointwise convolution (with Gated Linear Units), a depthwise 1-D convolution (kernel size 31), batch normalization, Swish activation, dropout, and residual addition:
   $$
   Z_i^{(c)} = Z_i^{(a)} + \text{Conv}(Z_i^{(a)})
   $$
5. **Second FFN & LayerNorm:** Another half-step FFN with Swish, followed by layer normalization:
   $$
   Z_i^{(o)} = \text{LayerNorm}\left(Z_i^{(c)} + \frac{1}{2} \cdot \text{FFN}(Z_i^{(c)})\right)
   $$
   
The conformer structure enables robust local feature extraction via convolution and global temporal context through self-attention, surpassing the modeling limits of CNN-, LSTM-, or transformer-only configurations [2502.11840].

## 4. Output Decoding and Temporal Smoothing

The output sequence (shape $T \times 256$) is linearly projected to six logit vectors $S^{(t,j)} \in \mathbb{R}^{M_j}$, one per chord component. Component-wise softmax yields probabilities:
$$
\beta_m^{(t,j)} = \frac{\exp S_m^{(t,j)}}{\sum_{m'} \exp S_{m'}^{(t,j)}}
$$
To enforce temporal smoothness, a chain-CRF decoder replaces naïve argmax decoding. The sequence probability is:
$$
P(Z|X) \propto \prod_{t=1}^T \phi(Z^{(t)}, X) \cdot \prod_{t=2}^T \psi(Z^{(t-1)}, Z^{(t)})
$$
where
$$
\phi(Z^{(t)}, X) = \exp \left( \sum_{j,m} I[m = z_j^{(t)}] \log \beta_m^{(t,j)} \right)
$$
$$
\psi(Z^{(t-1)}, Z^{(t)}) = \exp \left( -\gamma \cdot I[Z^{(t-1)} \neq Z^{(t)}] \right)
$$
This approach encourages label stability across adjacent frames, consistent with the piecewise-constant nature of chord sequences in music [2502.11840].

## 5. Reweighted Loss and Class Imbalance Mitigation

Given the substantial imbalance in chord-label frequencies, ChordFormer employs a reweighted cross-entropy loss:
$$
L = - \sum_{t=1}^T \sum_{j=1}^6 \sum_{m=1}^{M_j} w_m^{(j)} I[m=z_j^{(t)}] \log \beta_m^{(t,j)}
$$
Class weights are defined as:
$$
w_m^{(j)} = \min \left\{ \left( \frac{n_m^{(j)}}{\max_{m'} n_{m'}^{(j)}} \right)^{-\gamma}, w_{\max} \right\}
$$
where $n_m^{(j)}$ is the count of samples for class $m$ in component $j$, $\gamma \in [0,1]$ is a balancing exponent favoring rare classes, and $w_{\max}$ caps the maximum weight. Empirical results indicate optimal class-wise accuracy at $\gamma\approx0.5–0.7$, $w_{\max}\approx 10-20$. These settings amplify gradient contributions from rare chord types and elevate class-wise accuracy while minimally affecting frame-wise accuracy. Ablations confirm that the architecture is robust even under aggressive reweighting, preserving major/minor accuracy while boosting performance on rare extensions [2502.11840].

## 6. Training, Evaluation, and Comparative Analysis

Training utilizes the AdamW optimizer (initial learning rate $10^{-3}$, plateau scheduler), with early stopping once the learning rate drops below $10^{-6}$. Batches comprise randomly sampled 1 000-frame segments (≈23.2 s) from songs; mini-batch size is 24. Regularization includes dropout (rate $\approx$0.1 in all sublayers), batch normalization in convolution modules, and pre-norm residual connections.

On the Humphrey–Bello corpus (1 217 songs; 5-fold cross-validation), ChordFormer achieves:
- Frame-wise accuracy: 78.77%
- Class-wise accuracy: 38.84%
- MIREX score: 83.62%
In contrast, a CNN+BLSTM baseline yields 76.76%/33.15%/81.52%, respectively. Metric breakdowns indicate Root: 84.69%, Maj/Min: 84.09%, Triads: 77.55%, Sevenths: 72.28%. Confusion matrices show reduced misclassifications for extensions and rare chords, corresponding to effective local/global modeling and reweighted learning. Module ablations confirm that the conformer block delivers superior triad/extension recall over CNN-, transformer-, and BLSTM-based variants [2502.11840].

## 7. Architectural Significance and Empirical Insights

ChordFormer demonstrates that the fusion of convolution and self-attention within conformer blocks effectively unifies short- and long-range sequence modeling, addressing the core requirements of structural chord recognition. The structured decomposition of chord labels facilitates semantically meaningful parameter sharing and interpretable outputs. The introduction of a reweighted cross-entropy loss successfully mitigates performance degradation on low-frequency chord classes, a persistent issue in large-vocabulary regimes.

A plausible implication is that further advances in large-vocabulary symbolic music tasks may benefit from similar decompositions and hybrid modeling. The empirical gains in both frame- and class-wise accuracy establish ChordFormer as a reference architecture for robust and balanced audio chord recognition [2502.11840].

Source: https://www.emergentmind.com/topics/conformer-based-ultrasound-to-speech-conversion