Papers
Topics
Authors
Recent
Search
2000 character limit reached

XDementNET: MRI-Based Alzheimer’s Detection & Staging

Updated 11 July 2026
  • The paper introduces XDementNET, an MRI-based network that combines convolutional backbones with multi-residual and attention blocks for explainable Alzheimer’s staging.
  • It is evaluated across diverse datasets and classification setups, demonstrating near-perfect performance and improved clinical explainability over conventional saliency methods.
  • Its architecture leverages specialized spatial and grouped query attention to refine feature extraction, though it also highlights challenges like potential slice-level leakage and class imbalance.

Searching arXiv for the cited papers to ground the article. XDementNET is an explainable, attention-enhanced deep convolutional network for detecting and staging Alzheimer’s disease from MRI, introduced to address the clinical need for precise diagnosis together with transparent model rationales (Lincoln et al., 20 May 2025). It combines a convolutional backbone with multiresidual blocks, a specialized spatial attention block, Grouped Query Attention, and Multi-Head Attention, and is evaluated on Kaggle, OASIS-1, ADNI-1, and ADNI-2 under binary, multiclass, and cross-plane settings. In the same 2025 period, a distinct line of work proposed xEEGNet for explainable EEG-based dementia classification; the term “XDementNET” does not refer to that EEG model, which targets Alzheimer’s disease, frontotemporal dementia, and healthy controls from resting-state EEG rather than MRI (Zanola et al., 30 Apr 2025).

1. Nomenclature, scope, and clinical target

XDementNET is proposed for MRI-based Alzheimer’s disease detection and progression staging, with the stated motivation that MRI is widely used for Alzheimer’s disease because of its non-invasiveness, high tissue contrast, and spatial resolution (Lincoln et al., 20 May 2025). The paper frames the clinical problem around early and accurate identification of Alzheimer’s disease progression, and emphasizes that many high-performing AI systems remain “black boxes,” which can limit clinical acceptance because clinicians cannot verify which regions or features drive the prediction.

The network is designed for several classification regimes. On Kaggle and OASIS, it is used in a 4-class setup with NonDemented, VeryMildDemented, MildDemented, and Moderate Demented classes; in a 3-class setup where Mild and Moderate are merged; and in a binary demented-versus-nondemented setting. On ADNI-1, it is evaluated as a 3-class classifier across AD, MCI, and CN. On ADNI-2, it is evaluated in a 5-class setting with AD, CN, EMCI, MCI, and LMCI, and also in a 3-class setting where EMCI and LMCI are merged into MCI (Lincoln et al., 20 May 2025).

A necessary terminological distinction is that xEEGNet and XDementNET address related but different modalities and tasks. xEEGNet is expressly designed for explainable, compact EEG-based classification of Alzheimer’s disease, frontotemporal dementia, and controls using resting-state, eyes-closed EEG, whereas XDementNET is an MRI-based model focused on Alzheimer’s disease staging (Zanola et al., 30 Apr 2025). This suggests that the shared emphasis on explainability reflects a broader methodological trend in dementia-oriented medical AI rather than a single unified architecture.

2. Data sources, class structure, and preprocessing pipeline

XDementNET is assessed on four publicly accessible datasets, all converted into 2D image classification problems from MRI data (Lincoln et al., 20 May 2025). The Kaggle Alzheimer MRI 4 classes dataset contains 6,400 2D images derived from 200 subjects, with 32 axial slices per subject. OASIS-1 is described as cross-sectional T1-weighted MRI from 416 subjects, and the study slices the z-axis, selecting slices 100–160 and yielding 80,000 2D slice images per the paper’s count. ADNI-1 includes 643 subjects imaged with 1.5T MRI, with the middle 20 slices from axial, sagittal, and coronal planes converted into 2D images. ADNIGO/ADNI-2 uses only the axial plane, converting the middle 16 slices per subject into 2D images.

The reported class counts underscore substantial heterogeneity and imbalance across datasets. Kaggle contains CN 3,200, VMD 2,240, Mild 896, and AD 64 images. OASIS contains CN 67,222, VMD 13,725, Mild 5,002, and AD 488 images. ADNI-1 per plane contains CN 3,880, Mild/LMCI/MCI 6,320, and AD 2,660 images, with an all-plane combined total of 38,580. ADNI-2 contains CN 8,650, VMD 480, Mild/MCI 1,155, LMCI 144, and AD 8,346 images in a five-class setting (Lincoln et al., 20 May 2025).

The split methodology uses a two-stage split with random state 43. First, an 85/15 split reserves 15% for test; then 85/15 is applied on the remainder for train and validation. Augmentation is applied only to training data. The paper does not explicitly assert subject-level separation, and potential slice-level leakage is identified in the source description as a factor that should be considered when interpreting near-perfect results (Lincoln et al., 20 May 2025).

Preprocessing is described at the image level rather than at the level of neuroimaging harmonization. For ADNI datasets, 3D DICOM volumes of size 256×256×256256 \times 256 \times 256 are converted into middle slices and then to JPEG. For all datasets, images are converted to RGB, scaled to [0,1][0,1], and resized to 128×128×3128 \times 128 \times 3. A sharpening filter is applied that emphasizes the center pixel with value 5 and subtracts neighboring pixels with −1-1 to enhance edges. No skull-stripping, registration, or denoising is described for the presented pipeline (Lincoln et al., 20 May 2025).

The training set is further modified by augmentation and balancing. Rotation up to 15∘15^\circ, width and height shifts of 0.1, shear 0.2, zoom 0.2, and random horizontal flips are applied through TensorFlow ImageDataGenerator. Underrepresented classes are upsampled to match the second-largest class; if imbalance is extreme, the underrepresented class receives one-third as many augmented samples as the second-largest (Lincoln et al., 20 May 2025).

3. Architectural composition and computational pathway

XDementNET processes 128×128×3128 \times 128 \times 3 inputs through stacked Conv2D and MaxPooling blocks, followed by a Multi-Residual block, a specialized Spatial Attention block, Grouped Query Attention with Layer Normalization, Multi-Head Attention, Dropout, Global Average Pooling, and Dense classification heads (Lincoln et al., 20 May 2025). The paper positions the attention modules after convolutional feature extraction and before the final pooling and classification stages, so that attention modulates mid-level to high-level CNN features.

The multiresidual component is presented through the canonical residual mapping

y=F(x)+xy = F(x) + x

and, for the multibranch case,

y=(∑iFi(x))+xy = \left(\sum_i F_i(x)\right) + x

or

y=P([F1(x) ∣∣ F2(x) ∣∣ … ∣∣ Fm(x)])+x,y = P([F_1(x) \, || \, F_2(x) \, || \, \dots \, || \, F_m(x)]) + x,

with PP denoting a [0,1][0,1]0 projection when concatenation is used (Lincoln et al., 20 May 2025). The text does not enumerate explicit kernel sizes, strides, padding, or dilation for these branches, but the stated purpose is multi-scale feature extraction within a residual framework.

The specialized spatial attention block is introduced as a custom mechanism that computes a weight map [0,1][0,1]1 from feature maps [0,1][0,1]2 and applies it elementwise:

[0,1][0,1]3

The manuscript does not provide a closed-form equation beyond this qualitative formulation, nor does it specify pooling types, kernel sizes, or gating functions for the attention map generation (Lincoln et al., 20 May 2025).

The self-attention stack combines Grouped Query Attention and Multi-Head Attention. The core attention equation is given as

[0,1][0,1]4

with per-head projections

[0,1][0,1]5

and final composition

[0,1][0,1]6

Grouped Query Attention is described conceptually as refining feature extraction by modeling inter-feature relationships, but the paper does not specify the grouping formula, head count, or embedding dimensions (Lincoln et al., 20 May 2025).

Regularization and classification are handled through Dropout, Global Average Pooling, and Dense layers. The loss is categorical cross-entropy on one-hot labels,

[0,1][0,1]7

Although the paper states a design goal of lower computing costs and parameters, it does not report parameters or FLOPs quantitatively (Lincoln et al., 20 May 2025).

4. Training regimen, optimization choices, and reported performance

Training uses the Adam optimizer with initial learning rate [0,1][0,1]8, over 50 epochs, with batch size 8, on TensorFlow 2.0/Keras and Python 3.11 (Lincoln et al., 20 May 2025). Learning-rate control is performed through ReduceLROnPlateau with factor 0.7, patience 7, and minimum learning rate [0,1][0,1]9. The hardware configuration is NVIDIA RTX 3060 GPU, Intel Core i7 at 3.5 GHz, and 128 GB RAM.

The evaluation metrics reported are accuracy, precision, recall, F1-score, sensitivity, AUC-ROC, and RMSE. The equations are given in the manuscript as:

128×128×3128 \times 128 \times 30

128×128×3128 \times 128 \times 31

128×128×3128 \times 128 \times 32

128×128×3128 \times 128 \times 33

128×128×3128 \times 128 \times 34

128×128×3128 \times 128 \times 35

128×128×3128 \times 128 \times 36

128×128×3128 \times 128 \times 37

128×128×3128 \times 128 \times 38

The source notes that some equations have minor typographic bracket issues and reproduces them verbatim (Lincoln et al., 20 May 2025).

The headline performance claims are near-ceiling across datasets and task formulations. On Kaggle, XDementNET achieves 99.66% accuracy in 4-class classification, 99.63% in 3-class classification, and 100% in binary classification. On OASIS, it achieves 99.92%, 99.90%, and 99.95% respectively. On ADNI-1, it attains 99.08% accuracy for axial, 99.85% for sagittal, 99.50% for coronal, and 99.17% for all planes combined; the axial setup also reports AUC 99.80%. On ADNI-2, it reports 97.79% in the 5-class setup and 98.60% in the 3-class setup (Lincoln et al., 20 May 2025).

The source also records an inconsistency in the abstract, which lists “97.79% and 8.60%” for ADNI-2; the main results table and analysis clarify this as 98.60% for the 3-class setting, with the 8.60% figure identified as a typographical error (Lincoln et al., 20 May 2025). Confusion matrices are provided qualitatively, with errors described as rare and concentrated in adjacent disease stages, but exact per-class counts are not tabulated in the text.

5. Explainability framework and saliency comparison

A central claim of XDementNET is that explainability is not treated as an afterthought but is integrated through both architecture and post hoc saliency analysis (Lincoln et al., 20 May 2025). The paper compares the model’s attention maps with Grad-CAM, Score-CAM, Faster Score-CAM, and XGrad-CAM, arguing that the resulting visualizations are more concentrated and clinically plausible than those of the baseline explainability methods.

The manuscript provides the saliency equations verbatim. For Grad-CAM, it includes:

128×128×3128 \times 128 \times 39

−1-10

−1-11

−1-12

−1-13

−1-14

−1-15

−1-16

−1-17

For Score-CAM, it gives:

−1-18

−1-19

For XGrad-CAM, it gives:

15∘15^\circ0

15∘15^\circ1

The qualitative interpretation offered is that XDementNET’s saliency visualizations show concentrated red and yellow foci over diagnostically relevant brain regions, whereas some baselines produce broader or noisier activations (Lincoln et al., 20 May 2025). At the same time, the paper does not report quantitative localization metrics such as overlap with hippocampal ROIs. A plausible implication is that the explainability evidence is strongest at the level of comparative visualization rather than formal region-level validation.

6. Comparative position, limitations, and relation to adjacent explainable dementia models

The paper situates XDementNET against prior CNNs, transfer-learning models such as VGG16 and ResNet, EfficientNet, ViT-GRU hybrids, and attention-based 3D CNNs, and states that XDementNET matches or exceeds reported accuracies across Kaggle, OASIS, and ADNI benchmarks (Lincoln et al., 20 May 2025). Ablation-style experiments vary optimizers, learning rates, schedulers, preprocessing variants, and balancing or augmentation schemes; Adam with learning rate 15∘15^\circ2 and ReduceLROnPlateau are reported as the best-performing configuration. However, the paper does not isolate the quantitative contribution of each architectural component—multiresidual blocks, specialized spatial attention, GQA, and MHA—in separate module-level ablations.

The strongest methodological caveat concerns the possibility of overfitting or slice-level leakage. The source explicitly notes that the split strategy is fixed and consistent but does not explicitly guarantee subject-level separation, which is particularly salient because the data are generated as large numbers of 2D slices from comparatively smaller subject cohorts (Lincoln et al., 20 May 2025). Severe class imbalance is another reported risk, particularly on Kaggle where only 64 Moderate Demented slices are present. Merging classes can also simplify decision boundaries and inflate metrics.

Domain shift is identified as an additional concern, especially between ADNI-1 and ADNI-2, where cohort composition, imaging protocols, or processing pipelines may differ. The paper further notes that fairness across sex, age, or ethnicity is not addressed, and that clinical deployment would require cross-center validation, calibration, and uncertainty-aware evaluation (Lincoln et al., 20 May 2025).

In the broader landscape of explainable dementia AI, XDementNET is complementary rather than interchangeable with xEEGNet. xEEGNet demonstrates that a compact, fully interpretable EEG model with only 168 parameters can classify Alzheimer’s disease, frontotemporal dementia, and controls by explicitly filtering canonical EEG bands, learning band-specific scalp topographies, and classifying via band-power weights in decibels (Zanola et al., 30 Apr 2025). XDementNET, by contrast, addresses MRI-based Alzheimer’s disease staging through attention-augmented convolutional processing. This suggests a modality-specific divergence in explainability design: xEEGNet emphasizes explicit spectral and spatial decomposition, while XDementNET emphasizes attention-guided saliency over structural imaging features.

The reproducibility position of XDementNET remains partial. The datasets are public, the split random state is given as 43, and the preprocessing pipeline is described at a high level, but no public code repository or pretrained weights are reported (Lincoln et al., 20 May 2025). The recommendations for future work therefore focus on subject-level evaluation, multi-center external validation, multimodal integration with MRI, PET, CSF, and demographics, uncertainty quantification, graph-theoretic formulations, and human-in-the-loop explainability with quantitative localization metrics.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to XDementNET.