---
title: 'RML2018.01A Dataset: Dual-Use in MRI and AMR'
url: https://www.emergentmind.com/topics/rml2018-01a-dataset
type: topic
---

# RML2018.01A Dataset: Dual-Use in MRI and AMR

Searching arXiv for the cited papers and nearby usage of “RML2018.01A”.
Search query: 2102.07896
RML2018.01A is an identifier that refers to two distinct research datasets in contemporary literature. In speech-production MRI, it denotes a figshare corpus of 2D sagittal-view real-time magnetic resonance imaging (RT-MRI) videos with synchronized audio for 75 subjects, together with the corresponding raw multi-coil RT-MRI data, 3D volumetric vocal tract MRI during sustained speech sounds, and high-resolution static anatomical T2-weighted upper airway MRI [2102.07896]. In automatic modulation recognition (AMR), the same identifier denotes the DeepSig Inc. Over-the-Air corpus containing over 2.5 million complex I/Q examples of length 1024 across 24 modulation schemes, with channel impairments including multipath fading, carrier frequency offset, phase offset, and additive white Gaussian noise, over a nominal signal-to-noise ratio (SNR) range of $-20$ dB to $+30$ dB [2508.20193]. The shared name is therefore not a guarantee of shared domain, modality, or benchmark protocol.

## 1. Identifier disambiguation

In the cited literature, the designation RML2018.01A is used for two unrelated resources.

| Domain | Core contents | Scale |
|---|---|---|
| Speech-production MRI | 2D RT-MRI, synchronized audio, raw multi-coil RT-MRI, 3D volumetric MRI, static T2-weighted MRI | $N = 75$ healthy adults |
| Automatic modulation recognition | Complex I/Q time series over multiple modulation schemes and SNRs | Over 2.5 million examples |

This distinction is methodologically consequential. In the speech-production setting, RML2018.01A is organized around imaging rapidly moving articulators and dynamic airway shaping during speech, with explicit emphasis on raw acquisition data, reconstruction, artifact correction, and feature extraction [2102.07896]. In the AMR setting, RML2018.01A is a large over-the-air radio-frequency benchmark whose utility lies in modulation classification under channel impairments and varying SNR [2508.20193].

A common source of confusion is to assume that any citation to “RML2018.01A” refers to the radio-frequency benchmark. The speech-production corpus explicitly uses the same official identifier and version, RML2018.01A, with figshare DOI 10.6084/m9.figshare.13725546.v1. Conversely, the AMR literature summarized here uses “Original RML2018.01A (DeepSig Inc. Over-the-Air)” as the source dataset name. This suggests that domain context, accompanying modality description, and cited paper are necessary for correct interpretation.

## 2. Speech-production MRI corpus: subjects, tasks, and linguistic coverage

The speech-production instance of RML2018.01A contains data from $N = 75$ healthy adults, with gender distribution 40 F and 35 M, age range 18–59 years, 49 native American English speakers, and 26 non-native speakers with $L1 \neq$ English [2102.07896]. The represented first languages include Mandarin, Spanish, German, Korean, Telugu, Hindi, Vietnamese, Cantonese, Tamil, Portuguese, Greek, Russian, Gujarati, Twi, Bahasa Indonesia, and Native Hawaiian. All subjects performed American English materials, while the speaker $L1$s and accents provide a wide cross-section of global English pronunciation patterns.

The 2D RT-MRI component consists of scripted and spontaneous speech, approximately 17 min per subject. The scripted tasks include consonants in symmetric /VCV/ contexts (vcv1–vcv3, 30 s each, repeated twice), vowels in /bVt/ contexts (bvt, 30 s, repeated twice), four “shibboleth” phonologically rich sentences (30 s $\times 2$), and the Rainbow, Grandfather, and Northwind passages (30 s each, some repeated $\times 2$). The corpus also includes articulatory gestures and postures such as clench, open wide, yawn, swallow, a slow /i-a-u-i/ sequence, tongue-palate tracing, and singing “la” at highest and lowest pitch. The spontaneous component comprises descriptions of 5 pictures and answers to 5 open-ended questions, each 30 s.

The 3D volumetric MRI component lasts approximately 30 min and targets sustained speech sounds and postures. Sustained vowels are collected in /bVt/ contexts, including beet, bit, bait, bet, bat, pot, but, bought, boat, boot, put, bird, and abbot, each for 7 s and repeated twice. Sustained consonants in symmetric /VCV/ contexts include 13 stimuli, each 7 s and repeated twice, for example afa, ava, atha_thing, and ara. Sustained postures include breathe, clench, tongue-out, yawn, tip, and hold, again for 7 s and repeated twice. High-resolution static T2-weighted MRI at rest is acquired in axial, coronal, and sagittal sweeps.

The dataset was introduced to address the lack of open raw multi-coil RT-MRI data from an optimized speech production experimental setup. A common misconception is that speech-production MRI datasets of this type contain only reconstructed videos. Here, the corpus explicitly includes the corresponding first-ever public domain raw RT-MRI data in addition to reconstructed images and synchronized audio [2102.07896].

## 3. Speech-production MRI acquisition, reconstruction, and data organization

The 2D real-time MRI protocol uses a GE Signa Excite 1.5 T scanner (40 mT/m, 150 mT/m/ms slew) and a custom 8-channel upper airway array with 4 elements per cheek [2102.07896]. The sequence is a 13-interleaf spiral-out spoiled GRE with bit-reversed interleaf ordering, with $\mathrm{TR} = 6.004$ ms, $\mathrm{TE} = 0.8$ ms, flip angle $= 15^\circ$, $\mathrm{FOV} = 200 \times 200$ mm$^2$, slice thickness $= 6$ mm, in-plane resolution $= 2.4 \times 2.4$ mm$^2$ (84 $\times$ 84 pixels), and receiver bandwidth $= \pm 125$ kHz. The non-Cartesian spiral trajectory uses 13 arms for Nyquist, with 2 arms used per frame, yielding 83.28 frames/s. The stored data include multi-coil information for 8 channels, and k-space sampling density and locations are stored in the MRD header.

The 3D volumetric vocal tract MRI uses accelerated 3D GRE with Cartesian sparse Poisson-Disc sampling. The acquisition parameters are $\mathrm{TR} = 3.8$ ms, flip angle $= 5^\circ$, $\mathrm{FOV} = 200 \times 200 \times 100$ mm$^3$, matrix $= 160 \times 160 \times 80$, voxel size $= 1.25 \times 1.25 \times 1.25$ mm$^3$, and net acceleration $= 7\times$, with central $40 \times 20$ $k_y$-$k_z$ fully sampled for coil calibration and outer k-space Poisson-disc sampling. Scan time per stimulus is 7 s. High-resolution static anatomical T2-weighted MRI uses fast spin echo (FSE), with $\mathrm{TR} = 4600$ ms, $\mathrm{TE} = 120$–122 ms, echo train length $= 25$, in-plane $\mathrm{FOV} = 300 \times 300$ mm$^2$, in-plane resolution $= 0.5859 \times 0.5859$ mm$^2$, slice thickness $= 3$ mm, 29–70 slices per orientation, averages $= 1$, and scan time $\approx 3.5$ min per orientation.

The reference 2D RT-MRI reconstruction solves
$$
\min_m \|A m - d\|_2^2 + \lambda \|\nabla_t m\|_1,
$$
where $A$ encodes non-uniform FFT and coil sensitivities estimated via the Walsh method, $\nabla_t$ is the temporal finite-difference operator, and $\lambda = 0.08\,C$ with $C = \max$ intensity of the zero-filled reconstruction, selected via a sweep over $[0.008C \ldots 1C]$. The solver is nonlinear conjugate-gradient with Fletcher-Reeves and backtracking, with iteration limit $\leq 150$ or step $< 10^{-5}$. The reported reconstruction rate is $160.69 \pm 1.56$ ms/frame on Xeon E5-2640 v4 + Tesla P100 using MATLAB 2019b. The forward and inverse formulations are
$$
y = E x + \epsilon,
$$
$$
E = S F,
$$
and
$$
\hat x = \arg\min_x \|E x - y\|_2^2 + \lambda R(x).
$$
For 3D volumetric reconstruction, sparse-SENSE with isotropic spatial total variation is implemented in BART.

The corpus is approximately 966 GB and is publicly available via figshare at DOI 10.6084/m9.figshare.13725546.v1. The subject directory structure includes `2drt/raw/` for `*_raw.h5` MRD files, `2drt/recon/` for `*_recon.h5` HDF5 images, `2drt/audio/` for `*_audio.wav`, `2drt/video/` for `*_video.mp4`, `3d/recon/` for `*_recon.mat`, `3d/snapshot/` for `*_snapshot.png`, and `t2w/dicom/` for de-identified DICOM images. The naming convention is `<subID>_<modality>_<stimIndex>_<stimName>_r<rep>_<type>.*`. Metadata are distributed through `Subjects.xlsx`, `Stimuli.ppt`, and `metafile_public_<timestamp>.json`, the last of which records per-subject/task information, file existence, inspector notes, and visual and audio quality scores on 1–5 Likert scales for off-resonance blurring, video SNR, aliasing, and audio SNR. The dataset is free and public-domain under figshare terms, reconstruction code in MATLAB/Python is released under the MIT License, and no additional usage restrictions are specified beyond citation [2102.07896].

## 4. Radio-frequency modulation corpus: signal model and benchmark structure

In AMR, RML2018.01A denotes a radio-frequency dataset containing over 2.5 million complex I/Q examples, each a time series of length 1024, spanning 24 modulation schemes [2508.20193]. Channel impairments include multipath fading, carrier frequency offset, phase offset, and additive white Gaussian noise. The nominal SNR range is $-20$ dB to $+30$ dB.

The signal model is stated as
$$
r(n) = A(n)e^{j(\omega n + \theta)}x(n) + \sigma(n), \qquad n = 0,\ldots,N-1,
$$
with $\sigma(n) \sim \mathcal{CN}(0,\sigma^2)$. The in-phase and quadrature components are
$$
I(n) = \mathrm{Re}[r(n)], \qquad Q(n) = \mathrm{Im}[r(n)].
$$
The class-conditional hypothesis for class $k$ is
$$
H_k: x_k(n) = s_k(n) + \omega_k(n), \qquad \omega_k(n) \sim \mathcal{CN}(0,\sigma^2).
$$
The SNR definition is
$$
\mathrm{SNR} = 10 \cdot \log_{10}(P_s / P_n).
$$

The 2025 AMR study does not use the full dataset directly. Instead, it filters the original resource to 16 of the 24 modulations: PSK family (BPSK, QPSK, 8PSK), APSK family (16APSK, 32APSK, 64APSK, 128APSK), QAM family (16QAM, 32QAM, 64QAM, 128QAM, 256QAM), and other classes (AM-DSB-SC, AM-DSB-WC, FM, GMSK). The SNR range is further restricted to $-2$ dB through $+21$ dB. From each $(\text{class}, \text{SNR})$ pair, 1 000 examples are drawn, producing approximately 220 000 total signals. Although the original waveforms have length 1024, the Vision Transformer experiments reshape them into a $2 \times 512$ real-valued I/Q input.

A recurrent misconception is to read reported AMR accuracies as results on the original 24-class, full-SNR benchmark. The reported results summarized here are for the filtered 16-class subset over the restricted $-2$ dB to $+21$ dB range [2508.20193].

## 5. Limited-label protocol, augmentation, and evaluation in AMR

After construction of the filtered subset of approximately 220 k examples, the data are shuffled and split into 70% train (approximately 154 k), 10% validation (approximately 22 k), and 20% test (approximately 44 k) [2508.20193]. Each experiment is repeated five times with different random splits, and the reported values are averages.

The fine-tuning protocol uses limited-label regimes on the 70% training pool. Labeled fractions are 10%, 15%, and 20% of the training set, corresponding approximately to 15 k, 23 k, and 31 k labeled examples. The remaining training examples are treated as unlabeled. Unlabeled examples are pseudo-labeled at inference time if the model’s softmax confidence is at least 0.8, and these pseudo-labels enter the classification loss with weight 0.5.

At each training epoch, every signal receives exactly one transform with probability 1.0. The available transforms are rotation by $\phi \in \{0,\pi/2,\pi,3\pi/2\}$,
$$
(I + jQ)' = (I + jQ)e^{j\phi},
$$
horizontal flip,
$$
(I',Q') = (I,-Q),
$$
vertical flip,
$$
(I',Q') = (-I,Q),
$$
additive Gaussian noise,
$$
r'(n) = r(n) + n(n), \qquad n(n) \sim \mathcal{CN}(0,\sigma_n^2),
$$
scaling,
$$
r'(n) = \alpha r(n), \qquad \alpha \sim \mathcal{N}(1,\sigma_\alpha^2),
$$
time warping,
$$
r'(n) = r(n + \delta(n)),
$$
where $\delta(n)$ is a smooth random offset defined by a cubic spline, and magnitude warping, where $|r'(n)| = \mathrm{warp}(|r(n)|)$ via a random spline while phase is left intact.

The primary evaluation metric is classification accuracy,
$$
A = \frac{1}{N}\sum_{i=1}^N \mathbf{1}\{\hat y_i = y_i\}.
$$
Per-SNR accuracy is defined as
$$
A(\gamma) = \frac{\#\ \text{correct at SNR} = \gamma}{\#\ \text{total at SNR} = \gamma}.
$$
Overall performance across SNRs is reported both as a single overall accuracy and as curves of $A(\gamma)$ versus $\gamma$. The SNR-wise plots use 700 test samples per SNR level.

## 6. Reported performance, quality indicators, and significance

For the speech-production MRI corpus, the reported spatial resolution is 2.4 mm in 2D and 1.25 mm isotropic in 3D, with temporal resolution 83.28 fps for 2D RT-MRI [2102.07896]. Video SNR and audio SNR are rated from 1 (poor) to 5 (excellent) by an expert with 6 yr MRI experience. The common artifacts are off-resonance blurring near air-tissue boundaries and ringing/aliasing outside the field of view due to gradient non-linearity. The noise-temporal-smoothing tradeoff is illustrated for the reconstruction parameter sweep, and $\lambda = 0.08C$ is selected for optimal balance. The reported reconstruction speed of approximately 160 ms/frame is stated to allow near-real-time feedback. Demonstrated utility includes a quantitative speaking-rate histogram with mean $149.2 \pm 31.2$ wpm over read passages and observed inter-speaker variability in articulation timing.

For the AMR benchmark usage, the semi-supervised Vision Transformer with reconstruction-only pretraining reaches $68.21 \pm 2\%$, $71.01 \pm 2\%$, and $73.66 \pm 2\%$ test accuracy with 10%, 15%, and 20% labels, respectively [2508.20193]. Reconstruction plus contrastive pretraining yields $66.40 \pm 2\%$, $70.24 \pm 2\%$, and $71.41 \pm 2\%$, while contrastive-only pretraining yields $53.41 \pm 2\%$, $54.41 \pm 2\%$, and $56.41 \pm 2\%$. Under the 15% label setting, the reported overall accuracies are 61.84 for a fully supervised CNN, 78.50 for a fully supervised ResNet, 68.21 for a fully supervised ViT, and 71.01 for the semi-supervised ViT. In the reported comparison, the semi-supervised ViT with only 15% labels outperforms the fully supervised CNN by approximately 9 percentage points and the supervised ViT by approximately 2.8 percentage points, while approaching the ResNet. SNR-wise, the 15% label reconstruction-only model is approximately 30–40% overall at $-2$ dB, exceeds 70% by approximately 5 dB, and peaks near approximately 90% above 15 dB. Per-class highlights include 32APSK at 82.78% versus 80.44% for the supervised ResNet, AM-DSB-WC at 78.51% versus 78.12%, and FM and GMSK at approximately 100% in both cases.

Taken together, the two RML2018.01A usages occupy very different methodological roles. In speech science and MRI reconstruction, the identifier denotes a multispeaker corpus built around raw acquisition access, dynamic image reconstruction, artifact analysis, and anatomically grounded speech-production measurement [2102.07896]. In AMR, it denotes a large-scale I/Q benchmark supporting controlled studies of label efficiency, augmentation, and SNR robustness [2508.20193]. The principal interpretive requirement is therefore disambiguation: the same identifier names two unrelated datasets, and reported results, file structures, and evaluation protocols are comparable only within the appropriate domain context.

Source: https://www.emergentmind.com/topics/rml2018-01a-dataset