Papers
Topics
Authors
Recent
Search
2000 character limit reached

CleanCTG: Deep Learning for CTG Artefact Correction

Updated 8 July 2026
  • CleanCTG is a deep learning model designed for multi-artefact detection and reconstruction in cardiotocography, addressing errors like halving, doubling, and spikes.
  • It employs a dual-stage system combining multi-scale convolution, transformer encoders, and context-aware cross-attention to localize and correct artefacts.
  • The approach improves clinical analysis by reducing false alarms and enabling faster, more accurate fetal heart rate monitoring for timely intervention.

CleanCTG is a deep learning model for multi-artefact detection and reconstruction in cardiotocography (CTG), designed to address the frequent corruption of fetal heart rate (FHR) signals by halving and doubling errors, maternal heart rate contamination, missing segments, and spike artefacts. It is formulated as an end-to-end dual-stage system: a detection stage identifies multiple artefact types using multi-scale convolution and context-aware cross-attention, and a reconstruction stage applies artefact-specific correction branches and position-wise fusion to recover a cleaner FHR trace while preserving uncorrupted regions. The model was trained on over 800,000 minutes of physiologically realistic, synthetically corrupted CTGs derived from expert-verified clean recordings and was further evaluated on clinician-annotated clinical data and in preprocessing for Dawes–Redman™ analysis (Wong et al., 11 Aug 2025).

1. Clinical context and problem formulation

CTG is central to antepartum fetal monitoring, but FHR signals are frequently corrupted by diverse artefacts that obscure genuine physiological patterns such as accelerations, decelerations, and baseline behavior. The reported clinical consequences include false alarms, misdiagnosis, delayed intervention, and unnecessary cesareans. The paper situates this problem within a broader difficulty of CTG interpretation, noting poor inter-rater agreement, with κ≈0.12\kappa \approx 0.12–$0.39$, and persistently high false positives.

The artefact taxonomy used by CleanCTG comprises five classes. Halving error occurs when every other beat is dropped, halving the apparent rate and creating spurious decelerations or baseline drops that can mimic fetal bradycardia and mask variability. Doubling error records each beat twice, doubling the apparent rate and producing false accelerations or an elevated baseline that can mask decelerations or suggest tachycardia. Maternal heart rate contamination replaces the fetal signal with maternal rhythm, approximately $70$–$110$ bpm, obscuring fetal features and potentially mimicking a normal baseline. Missing segments arise from transducer displacement or fetal movement and disrupt trend continuity, baseline estimation, and rate variability measures. Spike artefacts are isolated abrupt changes of ±5\pm 5–$40$ bpm that create false accelerations or decelerations and distort short-term variability estimates.

The paper’s motivating claim is that conventional preprocessing is usually narrow in scope. Traditional approaches typically address only missing data through simple interpolation or rule-based filtering, and do not correct scaling errors, maternal overlap, spikes, or compound artefacts. Many deep-learning pipelines, by contrast, assume clean inputs or focus only on downstream classification. CleanCTG therefore advances a detect-then-correct formulation in which explicit artefact localization and artefact-specific reconstruction are treated as prerequisites for more reliable downstream interpretation.

2. System design and architectural principles

CleanCTG is described as a dual-stage, multi-branch, gated pipeline with dual-scale encoding (Wong et al., 11 Aug 2025). Its principal novelty lies in coupling a multilabel artefact detector to a reconstruction module that is conditional on the detected artefact profile, rather than using a single unified denoiser or a classification-only pipeline.

The detection stage processes a 1-minute local FHR segment together with its full 10-minute contextual window. Multi-scale convolutional feature extraction is applied separately to the local and contextual inputs, using multiple CNN layers with varying kernel sizes. The stated purpose is to preserve sharp local features in the 1-minute segment while learning broader temporal structure over the 10-minute context. These representations are then passed through separate transformer encoders with LayerNorm, multi-head self-attention, feedforward MLPs, and residual connections.

The central integration mechanism is context-aware cross-attention. All local tokens attend to all context tokens; the architecture does not restrict attention to a single CLS-style query. The attention operator is given as

Attention(Q,K,V)=softmax(QKTdk)V,\mathrm{Attention}(Q,K,V)=\mathrm{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V,

with Q=f1minWQQ=f_{\mathrm{1min}}W_Q, K=f10minWKK=f_{\mathrm{10min}}W_K, and V=f10minWVV=f_{\mathrm{10min}}W_V. The stated rationale is that broader temporal trends help disambiguate brief and extended corruptions and thereby strengthen artefact localization.

After cross-attention, class-specific attention pooling emphasizes timepoints most informative for each artefact type, and class-specific MLPs generate a 5-dimensional binary vector corresponding to halving, doubling, maternal heart rate contamination, missing segments, and spikes. The model therefore supports simultaneous multi-artefact detection within a single 1-minute segment, including compound scenarios such as maternal contamination sandwiched by missing segments.

This design also yields interpretable intermediate outputs. The paper reports per-class attention maps, per-class detection probabilities, and cross-attention maps showing local-context alignment. This suggests that the detector is intended not only as a gating mechanism for reconstruction but also as a diagnostic interface for auditing which regions and temporal dependencies support each artefact decision.

3. Reconstruction strategy and multi-task optimization

The reconstruction stage uses two levels of gating. First, global gates derived from detection probabilities activate only relevant reconstruction branches. Second, within an activated branch, attention mechanisms identify corrupted positions. This hierarchical gating is a defining property of CleanCTG: branch-specific computation is conditional on detected artefact type, and final signal synthesis is conditional on position.

For halving and doubling errors, CleanCTG does not rely on generic sequence denoisers. Instead, it uses mathematical correction branches. A multilayer attention mechanism produces position-wise binary masks $0.39$0 for class $0.39$1 and timepoint $0.39$2. At masked positions, amplitude scaling is applied directly: multiplication by $0.39$3 for halving and by $0.39$4 for doubling. This explicitly encodes the assumed corruption mechanism as a scaling error rather than a general missing-information problem.

For maternal heart rate contamination, missing segments, and spikes, CleanCTG uses transformer-based denoisers. Each such branch reconstructs the end-to-end signal and is modulated by a global detection gate:

$0.39$5

where $0.39$6 is the detection gate for class $0.39$7. The inclusion of the original signal in the reconstruction expression is important: it acts as a safety mechanism that reduces over-processing when the branch is not activated.

Final output generation is handled by position-wise attention fusion:

$0.39$8

where the branches $0.39$9 are $70$0original, halving, doubling, MHR, missing, spikes$70$1. The paper explicitly states that this layer chooses, at each timepoint, among branch outputs and the original signal, thereby preserving clean regions.

Training is conducted in two stages. The detection module is trained first with multilabel binary cross-entropy,

$70$2

with $70$3. After the detector reaches near-optimal AU-ROC, it is frozen. The reconstruction module is then trained with per-artefact MSE losses,

$70$4

aggregated as

$70$5

and combined with detection loss in

$70$6

No adversarial, perceptual, or smoothness terms are reported. Class imbalance is handled implicitly through the synthetic injection protocol and multilabel BCE.

4. Data construction and synthetic corruption protocol

The source data are drawn from the OxMat database (Oxford Maternity), spanning 1991–2024 (Wong et al., 11 Aug 2025). Original signals sampled at 4 Hz were down-sampled to 1 Hz. Analysis uses non-overlapping 10-minute windows for context, while detection and reconstruction operate on 1-minute segments with the full 10-minute parent window supplied to the detector.

The clean subset consists of 81,840 ten-minute segments that are corruption-free and expert-verified. From these, the study constructs a synthetic dataset of 818,400 one-minute segments by injecting five artefact types with realistic parameters and combinations. Corruption is limited to at most 50% per segment, and continuous noise per type is limited to at most 5% of segment length. Injection probabilities are reported as 5% each for halving and doubling, and 10% each for maternal heart rate contamination, missing segments, and spikes.

The resulting synthetic segment-level artefact proportions are 3.39% for halving, 3.40% for doubling, 3.28% for maternal heart rate contamination, 44.10% for missing segments, and 76.00% for spikes. For the external clinician-annotated set, corresponding proportions are 0.09%, 0.23%, 6.61%, 25.53%, and 3.97%, respectively. This distributional mismatch is one of the paper’s explicit indicators of domain shift.

External validation uses 1,019 ten-minute clinical CTGs independently annotated by clinicians and split into 10,190 one-minute segments. No extra preprocessing is applied, and labels specify the presence and type of artefact. For reconstruction and Dawes–Redman™ evaluation, the paper uses 933 sixty-minute clinical recordings, yielding 5,598 ten-minute segments; segments with more than 50% missing data within any one minute are excluded. The training-test split on synthetic data is 95% training, including validation, and 5% test. The optimizer is described only as adaptive, with modest learning rate and moderate batch size; hardware and runtime are not reported.

5. Empirical performance and ablation findings

On synthetic data, CleanCTG reports perfect artefact detection with AU-ROC = 1.00, sensitivity 99.90%, specificity 99.80%, and accuracy 99.50% (Wong et al., 11 Aug 2025). On the external clinical-annotated dataset of 10,190 minutes, it reports AU-ROC = 0.95, sensitivity = 83.44%, specificity = 94.22%, and accuracy = 88.83%. Spike detection is identified as the hardest case, with CleanCTG achieving AU-ROC 0.77 on spikes, approximately 14.93% improvement over MLP, and statistically significant superiority over comparators with $70$7. The comparator set comprises MLP, ResNet(1D), transformer classifier, bi-GRU, CCT, and TimesNet; the next-best overall clinical AU-ROC values are 0.92 for the transformer classifier, 0.90 for CCT, and 0.88 for TimesNet.

For reconstruction on synthetic corrupted segments, CleanCTG attains MSE = $70$8, outperforming Conv-Transformer Encoder at $70$9, TimesNet at $110$0, U-Net at $110$1, PatchTST at $110$2, linear interpolation at $110$3, and autoregression at $110$4. On clean-segment preservation, U-Net achieves the lowest reported MSE, $110$5, while CleanCTG reports $110$6, followed by Conv-Transformer at $110$7, TimesNet at $110$8, PatchTST at $110$9, and MLP AE at ±5\pm 50.

Selected per-artefact results are also reported. Doubling correction is near-perfect for CleanCTG, with MSE approximately ±5\pm 51. Maternal heart rate corruption is described as the hardest case; the reported synthetic-evaluation values are approximately 0.009 for CleanCTG, approximately 0.0016 for MLP, and approximately 0.0017 for U-Net. Error increases with corruption length across all models, but CleanCTG maintains the lowest MSE for typical clinical durations; conventional methods are competitive only when corruption is shorter than 3 seconds.

Ablation studies support the claim that artefact-specific branching and position-wise fusion are central to performance. Replacing halving and doubling mathematical branches with transformer branches increases MSE to ±5\pm 52. A unified single transformer yields ±5\pm 53, and five stacked shallow transformers yield ±5\pm 54. Replacing position-wise attention fusion with a last-layer MLP gives reconstruction MSE ±5\pm 55 and clean MSE ±5\pm 56; a last-layer LSTM yields ±5\pm 57 and clean MSE ±5\pm 58. The paper interprets these results as evidence that tailored correction branches and fine-grained fusion are critical both for denoising corrupted regions and for preserving clean segments.

6. Clinical integration, interpretability, and limitations

The paper evaluates CleanCTG as a preprocessing stage for the Dawes–Redman™ system on 933 clinical recordings (Wong et al., 11 Aug 2025). The protocol applies reconstruction to raw 60-minute CTGs and then runs standard Dawes–Redman™ analysis at 10 minutes and every 2 minutes thereafter until normality criteria are met. Specificity improves from 80.70% on raw traces to 82.70% on CleanCTG-denoised traces, while sensitivity is preserved, changing from 40.70% to 40.90%. Mean time to decision decreases from 22.8 minutes (21.08–24.52) to 20.6 minutes (18.94–22.18), a 9.64% improvement, and median time decreases from 18 minutes (IQR 10–32) to 12 minutes (10–32), a 33.33% reduction. The stated rationale is that removing spurious decelerations and accelerations and improving continuity allows normality criteria to be met earlier without sacrificing true positives.

The proposed clinical use is as a preprocessing step, with artefact detection outputs and branch-selection maps reviewed for interpretability. The interpretability mechanisms include artefact heatmaps from class-specific attention, per-class detection probabilities, visualizable position-wise masks for halving and doubling, and cross-attention maps that expose local-context alignment. A plausible implication is that the model is intended to support auditability at both the classification and reconstruction levels, rather than merely emitting a corrected trace.

Several limitations are explicit. The model depends heavily on synthetic training data, and the clinical-annotated set is limited relative to the synthetic corpus. Domain shift is observed, particularly in lower sensitivity for some artefacts, especially spikes. Inference speed and resource usage are not evaluated, although down-sampling to 1 Hz is presented as reducing computational overhead and favoring real-time potential. Coverage is limited to five major artefact classes; broader taxonomies such as saturation and baseline wander are identified as future extensions. Multi-center and multi-device validation is still required, and the paper suggests that adding uterine activity and maternal heart rate as explicit channels may improve robustness. Additional future directions include self-supervised pretraining, domain adaptation, threshold learning and calibration, and uncertainty modeling.

Reproducibility remains partial. The OxMat dataset is not publicly available due to privacy, and code and pretrained weights are not reported as publicly released. As a result, the principal enduring contribution of CleanCTG is methodological: it operationalizes a context-aware detect-then-correct strategy in which artefact localization, branch-specific correction, and position-wise preservation of clean signal are tightly coupled within a single CTG preprocessing framework.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to CleanCTG.