Papers
Topics
Authors
Recent
Search
2000 character limit reached

RML2018.01A Dataset: Dual-Use in MRI and AMR

Updated 9 July 2026
  • RML2018.01A is a dual-domain dataset that serves both speech-production MRI and automatic modulation recognition with distinct data organization and protocols.
  • In the MRI corpus, it provides 2D real-time and 3D volumetric imaging with synchronized audio from 75 subjects, enabling advanced articulatory analysis.
  • For AMR, it offers over 2.5 million complex I/Q examples across multiple modulation schemes, supporting robust studies under varied channel impairments.

Searching arXiv for the cited papers and nearby usage of “RML2018.01A”. Search query: (Lim et al., 2021) RML2018.01A is an identifier that refers to two distinct research datasets in contemporary literature. In speech-production MRI, it denotes a figshare corpus of 2D sagittal-view real-time magnetic resonance imaging (RT-MRI) videos with synchronized audio for 75 subjects, together with the corresponding raw multi-coil RT-MRI data, 3D volumetric vocal tract MRI during sustained speech sounds, and high-resolution static anatomical T2-weighted upper airway MRI (Lim et al., 2021). In automatic modulation recognition (AMR), the same identifier denotes the DeepSig Inc. Over-the-Air corpus containing over 2.5 million complex I/Q examples of length 1024 across 24 modulation schemes, with channel impairments including multipath fading, carrier frequency offset, phase offset, and additive white Gaussian noise, over a nominal signal-to-noise ratio (SNR) range of 20-20 dB to +30+30 dB (Ahmadi et al., 27 Aug 2025). The shared name is therefore not a guarantee of shared domain, modality, or benchmark protocol.

1. Identifier disambiguation

In the cited literature, the designation RML2018.01A is used for two unrelated resources.

Domain Core contents Scale
Speech-production MRI 2D RT-MRI, synchronized audio, raw multi-coil RT-MRI, 3D volumetric MRI, static T2-weighted MRI N=75N = 75 healthy adults
Automatic modulation recognition Complex I/Q time series over multiple modulation schemes and SNRs Over 2.5 million examples

This distinction is methodologically consequential. In the speech-production setting, RML2018.01A is organized around imaging rapidly moving articulators and dynamic airway shaping during speech, with explicit emphasis on raw acquisition data, reconstruction, artifact correction, and feature extraction (Lim et al., 2021). In the AMR setting, RML2018.01A is a large over-the-air radio-frequency benchmark whose utility lies in modulation classification under channel impairments and varying SNR (Ahmadi et al., 27 Aug 2025).

A common source of confusion is to assume that any citation to “RML2018.01A” refers to the radio-frequency benchmark. The speech-production corpus explicitly uses the same official identifier and version, RML2018.01A, with figshare DOI 10.6084/m9.figshare.13725546.v1. Conversely, the AMR literature summarized here uses “Original RML2018.01A (DeepSig Inc. Over-the-Air)” as the source dataset name. This suggests that domain context, accompanying modality description, and cited paper are necessary for correct interpretation.

2. Speech-production MRI corpus: subjects, tasks, and linguistic coverage

The speech-production instance of RML2018.01A contains data from N=75N = 75 healthy adults, with gender distribution 40 F and 35 M, age range 18–59 years, 49 native American English speakers, and 26 non-native speakers with L1L1 \neq English (Lim et al., 2021). The represented first languages include Mandarin, Spanish, German, Korean, Telugu, Hindi, Vietnamese, Cantonese, Tamil, Portuguese, Greek, Russian, Gujarati, Twi, Bahasa Indonesia, and Native Hawaiian. All subjects performed American English materials, while the speaker L1L1s and accents provide a wide cross-section of global English pronunciation patterns.

The 2D RT-MRI component consists of scripted and spontaneous speech, approximately 17 min per subject. The scripted tasks include consonants in symmetric /VCV/ contexts (vcv1–vcv3, 30 s each, repeated twice), vowels in /bVt/ contexts (bvt, 30 s, repeated twice), four “shibboleth” phonologically rich sentences (30 s ×2\times 2), and the Rainbow, Grandfather, and Northwind passages (30 s each, some repeated ×2\times 2). The corpus also includes articulatory gestures and postures such as clench, open wide, yawn, swallow, a slow /i-a-u-i/ sequence, tongue-palate tracing, and singing “la” at highest and lowest pitch. The spontaneous component comprises descriptions of 5 pictures and answers to 5 open-ended questions, each 30 s.

The 3D volumetric MRI component lasts approximately 30 min and targets sustained speech sounds and postures. Sustained vowels are collected in /bVt/ contexts, including beet, bit, bait, bet, bat, pot, but, bought, boat, boot, put, bird, and abbot, each for 7 s and repeated twice. Sustained consonants in symmetric /VCV/ contexts include 13 stimuli, each 7 s and repeated twice, for example afa, ava, atha_thing, and ara. Sustained postures include breathe, clench, tongue-out, yawn, tip, and hold, again for 7 s and repeated twice. High-resolution static T2-weighted MRI at rest is acquired in axial, coronal, and sagittal sweeps.

The dataset was introduced to address the lack of open raw multi-coil RT-MRI data from an optimized speech production experimental setup. A common misconception is that speech-production MRI datasets of this type contain only reconstructed videos. Here, the corpus explicitly includes the corresponding first-ever public domain raw RT-MRI data in addition to reconstructed images and synchronized audio (Lim et al., 2021).

3. Speech-production MRI acquisition, reconstruction, and data organization

The 2D real-time MRI protocol uses a GE Signa Excite 1.5 T scanner (40 mT/m, 150 mT/m/ms slew) and a custom 8-channel upper airway array with 4 elements per cheek (Lim et al., 2021). The sequence is a 13-interleaf spiral-out spoiled GRE with bit-reversed interleaf ordering, with TR=6.004\mathrm{TR} = 6.004 ms, TE=0.8\mathrm{TE} = 0.8 ms, flip angle +30+300, +30+301 mm+30+302, slice thickness +30+303 mm, in-plane resolution +30+304 mm+30+305 (84 +30+306 84 pixels), and receiver bandwidth +30+307 kHz. The non-Cartesian spiral trajectory uses 13 arms for Nyquist, with 2 arms used per frame, yielding 83.28 frames/s. The stored data include multi-coil information for 8 channels, and k-space sampling density and locations are stored in the MRD header.

The 3D volumetric vocal tract MRI uses accelerated 3D GRE with Cartesian sparse Poisson-Disc sampling. The acquisition parameters are +30+308 ms, flip angle +30+309, N=75N = 750 mmN=75N = 751, matrix N=75N = 752, voxel size N=75N = 753 mmN=75N = 754, and net acceleration N=75N = 755, with central N=75N = 756 N=75N = 757-N=75N = 758 fully sampled for coil calibration and outer k-space Poisson-disc sampling. Scan time per stimulus is 7 s. High-resolution static anatomical T2-weighted MRI uses fast spin echo (FSE), with N=75N = 759 ms, N=75N = 750–122 ms, echo train length N=75N = 751, in-plane N=75N = 752 mmN=75N = 753, in-plane resolution N=75N = 754 mmN=75N = 755, slice thickness N=75N = 756 mm, 29–70 slices per orientation, averages N=75N = 757, and scan time N=75N = 758 min per orientation.

The reference 2D RT-MRI reconstruction solves

N=75N = 759

where L1L1 \neq0 encodes non-uniform FFT and coil sensitivities estimated via the Walsh method, L1L1 \neq1 is the temporal finite-difference operator, and L1L1 \neq2 with L1L1 \neq3 intensity of the zero-filled reconstruction, selected via a sweep over L1L1 \neq4. The solver is nonlinear conjugate-gradient with Fletcher-Reeves and backtracking, with iteration limit L1L1 \neq5 or step L1L1 \neq6. The reported reconstruction rate is L1L1 \neq7 ms/frame on Xeon E5-2640 v4 + Tesla P100 using MATLAB 2019b. The forward and inverse formulations are

L1L1 \neq8

L1L1 \neq9

and

L1L10

For 3D volumetric reconstruction, sparse-SENSE with isotropic spatial total variation is implemented in BART.

The corpus is approximately 966 GB and is publicly available via figshare at DOI 10.6084/m9.figshare.13725546.v1. The subject directory structure includes 2drt/raw/ for *_raw.h5 MRD files, 2drt/recon/ for *_recon.h5 HDF5 images, 2drt/audio/ for *_audio.wav, 2drt/video/ for *_video.mp4, 3d/recon/ for *_recon.mat, 3d/snapshot/ for *_snapshot.png, and t2w/dicom/ for de-identified DICOM images. The naming convention is <subID>_<modality>_<stimIndex>_<stimName>_r<rep>_<type>.*. Metadata are distributed through Subjects.xlsx, Stimuli.ppt, and metafile_public_<timestamp>.json, the last of which records per-subject/task information, file existence, inspector notes, and visual and audio quality scores on 1–5 Likert scales for off-resonance blurring, video SNR, aliasing, and audio SNR. The dataset is free and public-domain under figshare terms, reconstruction code in MATLAB/Python is released under the MIT License, and no additional usage restrictions are specified beyond citation (Lim et al., 2021).

4. Radio-frequency modulation corpus: signal model and benchmark structure

In AMR, RML2018.01A denotes a radio-frequency dataset containing over 2.5 million complex I/Q examples, each a time series of length 1024, spanning 24 modulation schemes (Ahmadi et al., 27 Aug 2025). Channel impairments include multipath fading, carrier frequency offset, phase offset, and additive white Gaussian noise. The nominal SNR range is L1L11 dB to L1L12 dB.

The signal model is stated as

L1L13

with L1L14. The in-phase and quadrature components are

L1L15

The class-conditional hypothesis for class L1L16 is

L1L17

The SNR definition is

L1L18

The 2025 AMR study does not use the full dataset directly. Instead, it filters the original resource to 16 of the 24 modulations: PSK family (BPSK, QPSK, 8PSK), APSK family (16APSK, 32APSK, 64APSK, 128APSK), QAM family (16QAM, 32QAM, 64QAM, 128QAM, 256QAM), and other classes (AM-DSB-SC, AM-DSB-WC, FM, GMSK). The SNR range is further restricted to L1L19 dB through ×2\times 20 dB. From each ×2\times 21 pair, 1 000 examples are drawn, producing approximately 220 000 total signals. Although the original waveforms have length 1024, the Vision Transformer experiments reshape them into a ×2\times 22 real-valued I/Q input.

A recurrent misconception is to read reported AMR accuracies as results on the original 24-class, full-SNR benchmark. The reported results summarized here are for the filtered 16-class subset over the restricted ×2\times 23 dB to ×2\times 24 dB range (Ahmadi et al., 27 Aug 2025).

5. Limited-label protocol, augmentation, and evaluation in AMR

After construction of the filtered subset of approximately 220 k examples, the data are shuffled and split into 70% train (approximately 154 k), 10% validation (approximately 22 k), and 20% test (approximately 44 k) (Ahmadi et al., 27 Aug 2025). Each experiment is repeated five times with different random splits, and the reported values are averages.

The fine-tuning protocol uses limited-label regimes on the 70% training pool. Labeled fractions are 10%, 15%, and 20% of the training set, corresponding approximately to 15 k, 23 k, and 31 k labeled examples. The remaining training examples are treated as unlabeled. Unlabeled examples are pseudo-labeled at inference time if the model’s softmax confidence is at least 0.8, and these pseudo-labels enter the classification loss with weight 0.5.

At each training epoch, every signal receives exactly one transform with probability 1.0. The available transforms are rotation by ×2\times 25,

×2\times 26

horizontal flip,

×2\times 27

vertical flip,

×2\times 28

additive Gaussian noise,

×2\times 29

scaling,

×2\times 20

time warping,

×2\times 21

where ×2\times 22 is a smooth random offset defined by a cubic spline, and magnitude warping, where ×2\times 23 via a random spline while phase is left intact.

The primary evaluation metric is classification accuracy,

×2\times 24

Per-SNR accuracy is defined as

×2\times 25

Overall performance across SNRs is reported both as a single overall accuracy and as curves of ×2\times 26 versus ×2\times 27. The SNR-wise plots use 700 test samples per SNR level.

6. Reported performance, quality indicators, and significance

For the speech-production MRI corpus, the reported spatial resolution is 2.4 mm in 2D and 1.25 mm isotropic in 3D, with temporal resolution 83.28 fps for 2D RT-MRI (Lim et al., 2021). Video SNR and audio SNR are rated from 1 (poor) to 5 (excellent) by an expert with 6 yr MRI experience. The common artifacts are off-resonance blurring near air-tissue boundaries and ringing/aliasing outside the field of view due to gradient non-linearity. The noise-temporal-smoothing tradeoff is illustrated for the reconstruction parameter sweep, and ×2\times 28 is selected for optimal balance. The reported reconstruction speed of approximately 160 ms/frame is stated to allow near-real-time feedback. Demonstrated utility includes a quantitative speaking-rate histogram with mean ×2\times 29 wpm over read passages and observed inter-speaker variability in articulation timing.

For the AMR benchmark usage, the semi-supervised Vision Transformer with reconstruction-only pretraining reaches TR=6.004\mathrm{TR} = 6.0040, TR=6.004\mathrm{TR} = 6.0041, and TR=6.004\mathrm{TR} = 6.0042 test accuracy with 10%, 15%, and 20% labels, respectively (Ahmadi et al., 27 Aug 2025). Reconstruction plus contrastive pretraining yields TR=6.004\mathrm{TR} = 6.0043, TR=6.004\mathrm{TR} = 6.0044, and TR=6.004\mathrm{TR} = 6.0045, while contrastive-only pretraining yields TR=6.004\mathrm{TR} = 6.0046, TR=6.004\mathrm{TR} = 6.0047, and TR=6.004\mathrm{TR} = 6.0048. Under the 15% label setting, the reported overall accuracies are 61.84 for a fully supervised CNN, 78.50 for a fully supervised ResNet, 68.21 for a fully supervised ViT, and 71.01 for the semi-supervised ViT. In the reported comparison, the semi-supervised ViT with only 15% labels outperforms the fully supervised CNN by approximately 9 percentage points and the supervised ViT by approximately 2.8 percentage points, while approaching the ResNet. SNR-wise, the 15% label reconstruction-only model is approximately 30–40% overall at TR=6.004\mathrm{TR} = 6.0049 dB, exceeds 70% by approximately 5 dB, and peaks near approximately 90% above 15 dB. Per-class highlights include 32APSK at 82.78% versus 80.44% for the supervised ResNet, AM-DSB-WC at 78.51% versus 78.12%, and FM and GMSK at approximately 100% in both cases.

Taken together, the two RML2018.01A usages occupy very different methodological roles. In speech science and MRI reconstruction, the identifier denotes a multispeaker corpus built around raw acquisition access, dynamic image reconstruction, artifact analysis, and anatomically grounded speech-production measurement (Lim et al., 2021). In AMR, it denotes a large-scale I/Q benchmark supporting controlled studies of label efficiency, augmentation, and SNR robustness (Ahmadi et al., 27 Aug 2025). The principal interpretive requirement is therefore disambiguation: the same identifier names two unrelated datasets, and reported results, file structures, and evaluation protocols are comparable only within the appropriate domain context.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to RML2018.01A Dataset.