Papers
Topics
Authors
Recent
Search
2000 character limit reached

CTG-Insight: Multi-Agent CTG Analysis

Updated 7 July 2026
  • CTG-Insight is a multi-agent LLM framework that decomposes CTG signals into five medically defined features for clear, guideline-based interpretation.
  • It uses parallel feature agents to assess baseline, variability, accelerations, decelerations, and sinusoidal patterns, achieving state-of-the-art accuracy and transparent clinical decisions.
  • The architecture supports error localization and provides actionable, emotionally supportive explanations for both clinicians and expectant parents.

CTG-Insight is a multi-agent LLM framework for cardiotocography analysis that interprets fetal heart rate and uterine contraction traces through guideline-grounded feature decomposition rather than opaque end-to-end prediction. It was introduced to address a central problem in remote fetal monitoring: home-use systems commonly expose raw cardiotocography data without contextual explanation, leaving expectant parents with ambiguity and anxiety and leaving clinicians with limited interpretability from black-box models. The framework decomposes each 20-minute antepartum CTG trace into five medically defined features—baseline, variability, accelerations, decelerations, and sinusoidal pattern—and then synthesizes these feature-level judgments into an overall classification with a natural-language explanation (Sun et al., 29 Jul 2025).

1. Clinical setting and design rationale

CTG-Insight is situated in remote fetal monitoring, where cardiotocography consists of fetal heart rate (FHR) and uterine contractions (UC). The stated design objective is not only accurate classification, but also interpretability, clinical usefulness, and emotionally supportive explanation for expectant parents. In this setting, the difficulty is not the absence of data, but the absence of structured interpretation: home-use CTG apps may provide raw values or unannotated traces, while deep learning systems can achieve strong classification performance without exposing the reasoning that links signal morphology to fetal-health assessment.

The framework therefore adopts a clinically organized representation of CTG interpretation. Rather than directly predicting a binary label from raw time series, it mirrors guideline-based reasoning by decomposing the trace into medically defined components and assigning each one to a dedicated agent. This design is explicitly motivated by established medical guidelines and by the need for transparent, feature-level statements that can be aggregated into normal, suspicious, or pathological conclusions. A plausible implication is that CTG-Insight treats interpretability not as a post hoc explanation layer, but as the primary structure of inference.

The prompts encode criteria from the FIGO consensus guidelines and the German S1-guideline, adapted to 20-minute antepartum traces. For notation, the framework denotes the fetal heart rate by FHR(t)FHR(t) and the uterine contraction signal by UC(t)UC(t). The paper also notes that conventional signal analysis might estimate a baseline with a moving average over at least 10 minutes while excluding accelerations and decelerations, and might quantify variability through standard deviation or RMSSD; however, CTG-Insight does not operationalize these numerical metrics. Instead, the LLM agents apply guideline-based visual criteria directly to rendered CTG plots.

2. Guideline decomposition into five medically defined features

The five-feature decomposition is the core ontological structure of CTG-Insight. Each feature is defined using conventional obstetric terminology and classified into normal, suspicious, or pathological categories according to explicit thresholds (Sun et al., 29 Jul 2025).

Baseline is defined as the mean FHR maintained over at least 10 minutes in the absence of accelerations or decelerations, with the possibility of variation between successive 10-minute segments. The thresholds are normal at 110–160 bpm, suspicious at 100–109 bpm or 161–180 bpm, and pathological below 100 bpm or above 180 bpm. The paper additionally treats the baseline as suspicious if it cannot be determined because of persistent fluctuations such as sustained accelerations or decelerations.

Variability is defined as the degree of FHR change over time, evaluated as average bandwidth amplitude in 1-minute segments, with normal fluctuations occurring 3–5 times per minute. Normal variability requires both amplitude and frequency criteria: 5–25 bpm, 3–5 fluctuations per minute, present for most of the time. Suspicious variability is mildly abnormal without reaching pathological thresholds. Pathological variability is defined as less than 5 bpm lasting at least 15 minutes, or greater than 25 bpm lasting at least 10 minutes.

Accelerations are increases in FHR above baseline of more than 15 bpm lasting more than 15 seconds. Formally, an acceleration is a segment where FHR(t)−b(t)≥15FHR(t) - b(t) \ge 15 bpm and duration is at least 15 s. The thresholding scheme is unusual in that the absence of accelerations is categorized as pathological, two accelerations in 20 minutes is normal, and periodic occurrence with every contraction is suspicious.

Decelerations are decreases in FHR below baseline of more than 15 bpm lasting more than 15 seconds, formalized as b(t)−FHR(t)≥15b(t) - FHR(t) \ge 15 bpm with duration at least 15 s, with timing rules defined relative to UC(t)UC(t). The framework distinguishes early, variable, late, prolonged, and atypical variable decelerations. Early decelerations mirror uterine contractions and are usually benign; variable decelerations are abrupt, often cord-compression related; late decelerations begin more than 20 seconds after contraction onset and indicate uteroplacental insufficiency and possible hypoxia; prolonged decelerations persist for more than 3 minutes; atypical variable decelerations include loss of primary or secondary rise, slow return to baseline post-contraction, elevated baseline after contraction, biphasic form, loss of oscillation during deceleration, or resumption of baseline at a lower level. The thresholding is normal when no decelerations are present, suspicious for early, variable, or brief prolonged decelerations below 3 minutes, and pathological for late decelerations, prolonged decelerations persisting across more than two contractions or more than 3 minutes, and atypical variable decelerations.

Sinusoidal pattern is divided into true sinusoidal and pseudosinusoidal forms. A true sinusoidal pattern is a regular, smooth, sine-like signal with amplitude 5–15 bpm and frequency 3–5 cycles per minute, lasting at least 10 minutes and accompanied by absent accelerations. Pseudosinusoidal traces have saw-tooth-like morphology, are often transient for less than 10 minutes, and have normal patterns before or after. Thresholding is normal when no sinusoidal pattern is present, suspicious for pseudosinusoidal morphology lasting under 10 minutes, and pathological for a true sinusoidal pattern.

These criteria make the framework internally tri-valued even though the benchmark dataset is binary. This suggests a richer intermediate representational layer than the final evaluation labels provide, because suspicious cases are grouped with abnormal cases in the dataset.

3. Multi-agent architecture and inference workflow

CTG-Insight consists of five parallel Feature Agents—baseline, variability, accelerations, decelerations, and sinusoidal—and a final Aggregator Agent (Sun et al., 29 Jul 2025). Each Feature Agent receives a rendered CTG image with upper FHR and lower UC channels at a paper speed of 1 cm/min, together with a feature-specific prompt encoding the feature definition, decision rules, role, and formatting examples. The decelerations agent additionally receives the deceleration types and their timing criteria relative to uterine contractions.

The agents are instructed to classify their assigned feature into normal, suspicious, or pathological and to provide a concise explanation tied to observed patterns such as amplitude, duration, and timing versus contractions. Each feature output therefore contains both a structured label and a short rationale. The architecture is explicitly asynchronous at the feature-agent level: the five feature analyses run in parallel, and only after all are complete does the Aggregator Agent execute.

The Aggregator Agent consumes the five feature outputs and applies synthesis rules to produce the overall CTG classification. The aggregation mapping is deterministic at the label-combination level: the overall result is normal if all features are normal; suspicious if exactly one feature is suspicious and all others are normal; and pathological if at least one feature is pathological or at least two features are suspicious. The Aggregator then produces a natural-language explanation that cites which features drove the decision and how those feature judgments map to the guideline criteria.

The signal-processing pipeline is intentionally minimal. The implementation loads 20-minute antepartum segments sampled at 4 Hz, yielding 4,800 samples, and uses Python with Matplotlib to render CTG images for visual analysis. No additional resampling, filtering, artifact removal, or peak/trough detection is reported. Missing-data handling and explicit windowing heuristics are also not detailed. Consequently, feature quantification is not an explicit signal-processing stage but an implicit visual interpretation stage carried out by the LLM agents.

4. Experimental setup, baselines, and reported performance

The empirical evaluation uses the NeuroFetalNet dataset, which provides 20-minute CTG traces with FHR and UC sampled at 4 Hz and annotated by clinicians with binary labels: normal versus abnormal, with suspicious cases grouped as abnormal. Time-series deep learning experiments follow the dataset’s 9:1 train:test split. LLM-based experiments sample 50 test instances at random, consisting of 25 normal and 25 abnormal cases, and results are reported as averages over five repeated trials. The LLM agents are built with GPT-4.1, used in zero-shot mode without fine-tuning, and coordinated through parallel feature prompts and an overall synthesis prompt (Sun et al., 29 Jul 2025).

The comparison set includes a single-agent LLM baseline labeled Direct Prompt and three deep learning baselines: NeuroFetalNet, CNN+BiGRU, and ResNet. The reported metrics are Accuracy and F1-score. CTG-Insight achieves Accuracy 96.40% and F1 97.81%. The single-agent Direct Prompt baseline achieves Accuracy 79.80% and F1 80.10%. The deep learning baselines report Accuracy 94.23% and F1 94.20% for NeuroFetalNet, Accuracy 84.04% and F1 84.16% for CNN+BiGRU, and Accuracy 82.88% and F1 82.79% for ResNet.

The paper characterizes CTG-Insight as state of the art on the NeuroFetalNet evaluation subset. It also presents the single-agent system as an ablation of the decomposition strategy: all feature and overall prompts are concatenated into one instruction, yielding substantially lower performance. The stated hypothesis is that the single-agent baseline underperforms because long prompts dilute instructions, whereas the multi-agent formulation preserves focus on feature-specific rules.

Several elements commonly expected in a more complete benchmarking protocol are not reported. The paper does not provide per-class metrics, precision/recall, confusion matrices, confidence intervals, or statistical tests. Additional ablations, such as removal of specific feature agents or prompt variation studies, are also absent. Inter-rater reliability for the clinician labels is not reported.

5. Interpretability, output structure, and clinical significance

The novelty claimed for CTG-Insight lies in structured, guideline-grounded interpretability rather than in label prediction alone (Sun et al., 29 Jul 2025). A typical synthesized output consists of one explanation per feature followed by an overall decision. The illustrative example describes a normal baseline of approximately 140 bpm sustained across multiple 10-minute segments, normal variability of 10–20 bpm with 3–5 cycles per minute, two accelerations meeting the ≥15\ge 15 bpm and ≥15\ge 15 s criteria within 20 minutes, absence of decelerations meeting the ≥15\ge 15 bpm and ≥15\ge 15 s criteria and absence of late or prolonged events, and absence of a sinusoidal pattern, followed by an overall normal classification under FIGO/S1 synthesis rules.

This output form has two technical consequences. First, it externalizes the intermediate decision variables that conventional deep learning models usually keep latent. Second, it organizes the explanation around medically defined categories already used in obstetric interpretation. A plausible implication is that the framework supports error localization at the feature level: disagreements can be attributed to baseline, variability, acceleration, deceleration, or sinusoidal reasoning rather than only to an opaque end label.

The paper also frames interpretability in emotional as well as clinical terms. In remote monitoring, overcalling suspicious patterns may increase unnecessary anxiety or interventions, whereas mislabeling late or prolonged decelerations may delay escalation. Because the framework produces feature-level statements instead of raw traces or unqualified class outputs, it is designed to be more actionable for clinicians and more understandable for non-specialist users. At the same time, the work explicitly calls for clinician oversight, trust calibration, and human-centered validation rather than autonomous deployment.

6. Limitations, extensibility, and position within CTG-analysis research

The paper identifies several limitations of the current system (Sun et al., 29 Jul 2025). Validation is based on a small LLM evaluation subset of 50 images, so broader statistical assessment remains pending. Robustness to data shifts—such as different devices, sampling rates, intrapartum rather than antepartum recordings, noise, and artifacts—is unresolved. Borderline variability and atypical variable decelerations are highlighted as especially challenging for visual interpretation, and true versus pseudosinusoidal distinction is described as clinically difficult. The framework also excludes contextual metadata such as gestational age and labor stage, even though these factors can materially affect interpretation.

Methodologically, CTG-Insight does not perform explicit numeric feature extraction, uncertainty quantification, or calibration. Runtime, resource use, token costs, code availability, seeds, and hyperparameters are not reported. Real-time streaming and state tracking are identified as requirements for deployment, and future work is proposed in real-time operation, mobile or wearable integration, stakeholder-centered evaluation, robustness to noise and data shift, and uncertainty estimation.

Its research position is clearer when placed beside other CTG-analysis traditions. Prior work on the CTU-UHB dataset used a modified GANomaly anomaly detector trained primarily on normal FHR and reported F1 =0.752=0.752 and balanced accuracy UC(t)UC(t)0, with a focus on semi-supervised abnormality detection rather than transparent, guideline-level interpretation (Bertieaux et al., 2022). Another line of work used segmented FHR windows with a one-dimensional convolutional neural network and reported sensitivity UC(t)UC(t)1, specificity UC(t)UC(t)2, and AUROC UC(t)UC(t)3 for abnormal birth-outcome detection at a 200-sample window size, emphasizing automatic feature learning from balanced windows rather than explanatory decomposition (Fergus et al., 2019). A third approach engineered CTG and clinical features, including ARMA-derived measures of FHR response to contractions, and reported AUC UC(t)UC(t)4, TPR UC(t)UC(t)5, and FPR UC(t)UC(t)6 when CTG features were combined with EHR variables and labor-stage duration (O'Sullivan et al., 2021). Against that background, CTG-Insight is distinctive in treating interpretability as the organizing principle of the model itself: it decomposes CTG into baseline, variability, accelerations, decelerations, and sinusoidal pattern, aligns each decision with FIGO and S1 criteria, and only then synthesizes an overall label.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to CTG-Insight.