Evaluating the Generalization of Neuroimaging Foundation Models on African Brain MRI
Abstract: Neuroimaging foundation models pretrained on large, predominantly western cohorts are increasingly proposed as general-purpose backbones for brain MRI analysis. Yet, their ability to generalize to underrepresented clinical populations remains largely untested. We evaluate four recent foundation models (BrainIAC, Neuro-JEPA, NeuroVFM, and Primus) on a three-way diagnostic classification task (Control, Dementia, Parkinson's disease) using a cohort of 88 subjects from a Nigerian clinical brain MRI dataset, across four modality configurations (T1w, T2w, T1w+T2w, FLAIR), and compare against an end-to-end trained ViT3D baseline. The frozen backbones collapse to majority-class predictions, while Neuro-JEPA on FLAIR shows modest but still limited discrimination. In contrast, the end-to-end trained ViT3D achieves higher accuracy and MCC on every task (up to 53.4% accuracy, MCC=0.27) and is the only model with non-trivial recall. Our findings suggest that these frozen neuroimaging foundation models are insufficient for fine-grained diagnostic classification in small, non-western clinical cohorts, motivating parameter-efficient adaptation and broader multi-site external validation for equitable deployment in global health settings.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper studies whether advanced artificial-intelligence models can correctly analyze brain MRI scans from African patients.
These AI models are called neuroimaging foundation models. They are trained beforehand using very large collections of brain scans, mostly from Western countries. The researchers wanted to know whether these models also work well on MRI scans collected in Nigeria.
This is an important question because medical technology should work fairly well for people from different countries and backgrounds—not only for the groups included in the original training data.
2. What questions did the researchers ask?
The main research questions were:
- Can four existing brain MRI foundation models tell the difference between:
- healthy people,
- people with dementia, and
- people with Parkinson’s disease?
- Do these models work well when they are used on Nigerian brain scans?
- Which type of MRI scan works best:
T1,T2,FLAIR, or a combination ofT1 + T2? - Do the foundation models perform better than a simpler AI model trained directly on the Nigerian data?
- Which parts of the brain do the models use when making their decisions?
3. How was the research carried out?
The data
The researchers used MRI data from 88 people in Nigeria:
- 31 people with dementia
- 22 people with Parkinson’s disease
- 35 healthy controls
The scans came from three medical centers in Nigeria. The MRI machines were not all the same; they had different strengths and produced images with different levels of quality. Some people were also missing one or more types of MRI scans.
This made the dataset more like the conditions found in real hospitals, but also made the problem more difficult for the AI models.
The MRI types
The study examined several kinds of MRI scans:
- T1-weighted scans, which show brain structure clearly
- T2-weighted scans, which can show different types of tissue and fluid
- FLAIR scans, which are especially useful for highlighting certain brain abnormalities
- T1 + T2 scans, which combine two types of information
Before testing the models, the researchers cleaned the images and removed non-brain parts, such as the skull. This is similar to cropping a photograph so that the computer focuses only on the important object.
The AI models
The researchers tested four pretrained foundation models:
- BrainIAC
- Neuro-JEPA
- NeuroVFM
- Primus
These models had already learned general patterns from many brain scans. The researchers froze the main parts of these models, meaning they did not allow the models to change their learned knowledge. They only added a small final layer that guessed whether a scan showed dementia, Parkinson’s disease, or a healthy brain.
This is like giving a student a textbook they are not allowed to rewrite, then asking them to use it to answer questions about a new subject.
The researchers also tested a ViT3D model. Unlike the foundation models, this model was trained from the beginning using the Nigerian data. A ViT3D is a type of computer-vision model designed to study three-dimensional objects, such as MRI volumes.
Measuring performance
The researchers used five-fold cross-validation. This means they divided the data into five parts, trained the model on four parts, and tested it on the remaining part. They repeated this process five times so that every part was used for testing.
They measured performance in several ways:
- Accuracy: the percentage of correct answers
- Balanced accuracy: accuracy that gives equal importance to each group
- F1 score: a measure of how well the model identifies each category
- MCC: a score that checks whether predictions are better than random guessing, especially when groups are different sizes
- AUC: how well the model separates one group from the others across many possible decision settings
The researchers also used Grad-CAM, an explanation method that creates a heat map showing which parts of an MRI image influenced the model’s decision. It is similar to asking an AI, “Which parts of the picture did you look at most closely?”
4. What did the researchers find?
The frozen foundation models performed poorly overall
Most of the foundation models struggled to distinguish the three groups. Several models mostly predicted the same category—usually dementia—rather than making balanced predictions.
For example, some models correctly identified dementia patients but almost never correctly identified people with Parkinson’s disease or healthy controls.
This means that a model can appear to perform reasonably well if one group is common, even though it is failing badly for the other groups.
The model trained on Nigerian data performed better
The ViT3D model, which was trained directly on the Nigerian scans, performed better than the frozen foundation models on nearly every task.
Its best results came from combining T1 and T2 scans:
| Measure | ViT3D result |
|---|---|
| Accuracy | 53.3% |
| Balanced accuracy | 48.2% |
| Macro-F1 | 43.4% |
| MCC | 0.27 |
| AUC | 0.672 |
These results are still not good enough for use as a dependable medical diagnostic tool. However, they were better than the other models tested.
The ViT3D model was also the only model that showed meaningful recognition of all three groups. For the combined T1 + T2 scans, it correctly identified:
- 11 of 22 healthy controls
- 26 of 31 people with dementia
- 8 of 22 people with Parkinson’s disease
Neuro-JEPA showed some promise
Neuro-JEPA was the strongest of the frozen foundation models, especially when using FLAIR scans.
On FLAIR images, it achieved:
- 44.8% accuracy
- 40.9% balanced accuracy
- 0.128 MCC
- 0.684 AUC
This suggests that FLAIR scans may contain useful information for distinguishing these conditions, and that Neuro-JEPA may have learned some helpful patterns. However, its performance was still limited and not strong enough for reliable clinical use.
The models may have looked at the wrong places
The explanation maps showed that some models did not focus clearly on important brain structures.
- BrainIAC and NeuroVFM often spread their attention across much of the image.
- Primus sometimes focused on the outer edges of the brain.
- ViT3D often focused on one side or edge of the image.
This raises the possibility that the models learned shortcut features. A shortcut feature is an irrelevant clue that helps the model guess the answer without understanding the real medical cause.
For example, the model might accidentally use differences caused by the scanner, image position, or hospital rather than signs of dementia or Parkinson’s disease. The researchers could not prove that this happened, but the image explanations suggest that it may have.
The results were uncertain because the dataset was small
The study used fewer than 100 people. With such a small dataset, performance can change a lot depending on which patients are included in the training and testing groups.
The researchers found statistically significant advantages for ViT3D on some T1 comparisons, but not for most other comparisons. This means that larger studies are needed before making strong conclusions.
5. Why are these findings important?
The study challenges the idea that a powerful AI model trained on huge datasets will automatically work well everywhere.
A model trained mainly on Western brain scans may struggle with Nigerian scans because of differences in:
- scanner types and image quality
- hospital procedures
- patient populations
- disease patterns
- missing MRI sequences
- age and demographic characteristics
In other words, the model may have learned to recognize the training data rather than the general medical features needed in every country.
The findings show that local testing is essential. Before an AI system is used in hospitals, it should be checked on data from the people and machines it will actually encounter.
6. What could happen next?
Future research could improve these models by:
- training them with more African and other underrepresented populations
- testing them at many hospitals in different countries
- allowing the pretrained models to adapt slightly to local data
- using methods such as LoRA, which changes only a small number of model parameters instead of retraining everything
- combining more MRI types, such as
T1 + T2 + FLAIR - checking carefully for shortcut learning
- collecting larger and more balanced datasets
Conclusion
The paper shows that current frozen brain MRI foundation models do not automatically generalize well to a small Nigerian clinical dataset. Most of them made unbalanced predictions and often failed to recognize Parkinson’s disease and healthy controls.
A model trained directly on the Nigerian scans performed better, although its results were still not accurate enough for diagnosis. Neuro-JEPA performed somewhat better on FLAIR images, but it also had important limitations.
The main lesson is that medical AI must be tested and adapted for different populations. Building fair and useful healthcare technology will require more diverse brain-imaging data, especially from Africa and other regions that have been underrepresented in AI research.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- The findings are based on only 88 participants from Nigeria, limiting statistical power and making the reported performance estimates highly uncertain, especially for Parkinson’s disease with 22 subjects.
- The study lacks independent external validation on datasets from other Nigerian sites, African countries, scanners, and clinical centers, so it remains unclear whether the observed failures are specific to this cohort or generalize across African populations.
- The effects of site, scanner manufacturer, field strength, acquisition protocol, and image quality are not disentangled. The dataset combines scanners ranging from 0.3 T to 1.5 T, but performance is not stratified by these factors.
- Demographic and clinical confounding is not adequately controlled. The control group is substantially younger than the dementia and Parkinson’s groups, and the analyses do not establish whether models exploit age-related or other non-pathological differences rather than disease-specific anatomy.
- The diagnostic labels and clinical characterization are insufficiently detailed. The paper does not report diagnostic criteria, disease severity, medication status, disease duration, comorbidities, or whether diagnoses were confirmed using standardized procedures.
- The cohort may not represent the broader African population. Participants come from three Nigerian centers, with no analysis of ethnicity, geographic diversity, socioeconomic factors, recruitment setting, or rural versus urban populations.
- The impact of missing MRI sequences is unresolved. Different modality-specific experiments use different participant subsets, but the paper does not assess whether missingness is systematic or whether it biases comparisons across modalities.
- The acquisition-selection procedure may introduce evaluation bias. The lowest-BRISQUE acquisition was selected, but the paper does not clarify whether this selection was performed independently within each cross-validation split or whether BRISQUE is clinically valid for choosing among repeated brain MRI acquisitions.
- Preprocessing consistency across models is unclear. The paper states that model-specific preprocessing was applied but does not fully specify or experimentally compare registration, resampling, intensity normalization, orientation handling, spatial cropping, and resolution differences.
- The frozen-backbone protocol does not determine whether the foundation models themselves are unsuitable or whether they simply require adaptation. No systematic comparison is made between frozen probing, full fine-tuning, partial fine-tuning, adapters, LoRA, prompt-based adaptation, or other parameter-efficient methods.
- The linear probes are not compared with stronger non-linear downstream heads. It remains unknown whether multilayer perceptrons, attention pooling, convolutional aggregation, or class-balanced classifiers could extract more useful information from the frozen patch representations.
- The comparison with ViT3D is not fully controlled for training conditions and model capacity. The paper does not provide a systematic parameter-count, pretraining-initialization, augmentation, hyperparameter, or compute-budget comparison between the end-to-end baseline and the frozen models.
- The reported superiority of ViT3D may reflect overfitting or shortcut learning. Its apparent advantage is based on a very small sample, and the study does not test whether performance persists under repeated cross-validation, nested model selection, or independent test evaluation.
- Hyperparameter-selection procedures are not described in sufficient detail. It is unclear how dropout, learning rate, stopping criteria, augmentation, class weighting, and other choices were selected without information leakage across folds.
- The statistical analysis is underpowered and potentially inappropriate for the experimental design. Paired -tests on only five fold-level observations provide limited evidence, and the study does not report confidence intervals, corrected multiple comparisons, permutation tests, bootstrap uncertainty, or effect sizes.
- No confidence intervals or subject-level variability estimates are reported for the performance metrics. Mean five-fold scores alone do not show how unstable the models are across splits or subjects.
- Class imbalance is only partially addressed. Although balanced accuracy and MCC are reported, the training loss does not appear to use class weighting, focal loss, resampling, or other imbalance-aware strategies.
- The study does not evaluate calibration or clinical decision utility. Reliability, expected calibration error, Brier score, decision-curve analysis, and clinically meaningful operating points are absent.
- The multi-class formulation may not reflect the clinical diagnostic problem. The analysis does not examine clinically relevant binary tasks, such as dementia versus control or Parkinson’s disease versus control, nor does it assess hierarchical or differential-diagnosis settings.
- The biological and clinical meaning of the learned representations is not established. Performance metrics do not reveal whether the models encode neurodegenerative pathology, demographic attributes, scanner signatures, or acquisition artifacts.
- The Grad-CAM analysis is exploratory and lacks quantitative validation. No anatomical region-of-interest comparison, localization metric, expert assessment, perturbation test, occlusion experiment, or faithfulness evaluation is provided.
- The suspected shortcut feature in ViT3D is not experimentally tested. The paper does not determine whether the edge activation reflects padding, cropping, skull-stripping artifacts, scanner borders, site markers, motion, coil effects, or other acquisition-specific signals.
- Attribution maps are averaged across subjects and normalized, potentially obscuring individual-level behavior. The study does not report attribution variability, statistical comparisons between classes, or whether the observed patterns are consistent for correctly and incorrectly classified subjects.
- The modality-specific Grad-CAM procedure for multimodal inputs is not validated. Masking other modalities to zero may create inputs outside the model’s training distribution, making the resulting modality attributions difficult to interpret.
- The value of richer multimodal combinations remains unknown. Although the paper evaluates T1w+T2w, it does not test T1w+T2w+FLAIR, alternative fusion strategies, missing-modality methods, or modality-specific adaptation.
- The effect of image quality and motion artifacts is not quantified. BRISQUE is used for acquisition selection, but the study does not evaluate performance as a function of image quality or test robustness to realistic degradation.
- The contribution of pretraining-domain mismatch is not isolated. The models differ in architecture, objective, pretraining dataset, modality composition, and imaging source, so the study cannot determine whether failures arise primarily from geographic population shift, scanner shift, modality mismatch, objective design, or architecture.
- The role of pretraining data diversity is not examined. The paper does not compare models trained with African or geographically diverse data against models trained predominantly on Western cohorts under matched protocols.
- Potential demographic or sensitive-attribute leakage is not assessed. The models are not tested for encoding age, sex, site, scanner, or other attributes that could drive predictions and undermine equitable deployment.
- No fairness analysis is reported across demographic or clinical subgroups. Performance differences by sex, age range, site, field strength, disease severity, or other relevant strata remain unknown.
- The study does not evaluate robustness to distribution shifts likely in routine African clinical deployment. There are no experiments involving unseen scanners, altered protocols, lower resolution, missing sequences, motion, intensity variation, or changes in disease prevalence.
- The practical deployment requirements are not quantified. Inference time, memory use, hardware requirements, preprocessing burden, annotation costs, and the cost of local calibration or adaptation are not reported.
- The models’ reproducibility and implementation sensitivity are unclear. The paper does not report random-seed variability, exact software versions, checkpoint details, preprocessing code, or whether all published checkpoints can be reproduced under the stated protocol.
- The study does not establish whether the models are useful for clinically meaningful outcomes beyond diagnosis. Prognosis, disease staging, progression prediction, treatment response, and screening applications are left unexplored.
- The conclusion that African disease etiology profiles contribute to poor transfer is speculative. The experiments do not measure or compare etiological, phenotypic, genetic, or comorbidity differences between the Nigerian cohort and the pretraining populations.
- The optimal scale and composition of local adaptation data remain unknown. The paper does not estimate how many labeled or unlabeled African scans are required to achieve reliable gains through domain adaptation or fine-tuning.
- The impact of self-supervised pretraining on African MRI is not tested directly. It remains unresolved whether pretraining on unlabeled local or regionally diverse scans would improve transfer more effectively than adapting the existing foundation models.**
Practical Applications
Immediate Applications
The study does not support immediate autonomous diagnosis or patient-facing deployment: frozen foundation models frequently collapsed toward the majority class, and even the best reported model achieved only moderate performance on a small, heterogeneous cohort. The most defensible near-term applications are therefore research, quality assurance, and carefully supervised clinical-support workflows.
- African neuroimaging benchmark for model evaluation — Academia and AI development
- Use the Nigerian cohort and its modality-specific splits as an external validation benchmark for brain MRI foundation models, conventional deep-learning models, and domain-adaptation methods.
- Compare models using balanced accuracy, macro-F1, MCC, per-class recall, and AUC rather than accuracy alone, since majority-class prediction can produce misleadingly acceptable accuracy.
- Potential tool: a reproducible evaluation package containing standardized preprocessing, subject-level cross-validation, modality availability checks, and class-wise reporting.
- Dependencies: access to the dataset, appropriate data-governance approvals, larger multi-site cohorts, and prevention of subject or acquisition leakage.
- Baseline selection for new African neuroimaging AI projects — Academia, hospitals, and software vendors
- Treat an end-to-end 3D ViT or other locally trained model as a required comparator instead of assuming that a frozen foundation model is superior.
- For small datasets, teams can reproduce the paper’s workflow: HD-BET skull stripping, acquisition-quality selection using BRISQUE, modality-specific preprocessing, and subject-level cross-validation.
- Practical outcome: a model-development checklist that requires comparison against simple supervised baselines before investing in foundation-model integration.
- Dependencies: sufficient labeled cases, careful regularization, and external validation to determine whether apparent gains come from pathology or dataset-specific artifacts.
- MRI data-quality control and acquisition auditing — Hospitals and radiology operations
- Use the paper’s preprocessing components as an initial quality-control workflow:
- automated skull stripping with tools such as
HD-BET; - no-reference image-quality scoring with BRISQUE;
- identification of missing T1w, T2w, and FLAIR sequences;
- scanner and field-strength tracking.
- This can help radiology departments flag low-quality scans, inconsistent acquisitions, or cases requiring manual review before AI analysis.
- Dependencies: BRISQUE scores are not equivalent to diagnostic image quality, and thresholds must be calibrated for low-field scanners, local protocols, and different anatomical sequences.
- Human-in-the-loop research triage for neurodegenerative imaging — Healthcare
- A locally trained model could be used to prioritize scans for expert review or research recruitment, particularly for dementia-related cases where the study found relatively stronger class-wise performance than for Parkinson’s disease.
- Outputs should be presented as risk scores or review priorities, not definitive diagnoses.
- Dependencies: prospective validation, calibrated probabilities, meaningful sensitivity for Parkinson’s disease and control cases, clinician oversight, and clear escalation procedures. The reported performance is insufficient for unsupervised clinical triage.
- Multimodal sequence prioritization — Radiology workflow design
- The relatively stronger performance of Neuro-JEPA on FLAIR and of ViT3D on T1w+T2w can inform experiments on which sequences provide the greatest incremental value when scan time, cost, or availability is limited.
- A hospital could evaluate whether acquiring T1w and T2w together, or prioritizing FLAIR in selected protocols, improves research classification or clinical review efficiency.
- Dependencies: the result is exploratory and based on small, non-equivalent modality subsets; it should not be interpreted as evidence that one sequence is clinically sufficient.
- Model-audit and shortcut-detection workflow — Responsible AI and medical software
- Use Grad-CAM or related attribution methods to identify diffuse attention, peripheral activation, or repeated hotspots that may indicate scanner borders, padding, registration artifacts, or other shortcuts.
- Such maps can be incorporated into a pre-deployment audit dashboard alongside performance by scanner, site, modality, age, and diagnosis.
- Dependencies: attribution maps are diagnostic aids rather than causal explanations. Suspected shortcuts must be tested through masking, perturbation, scanner-stratified validation, and prospective data collection.
- Clinical data infrastructure and African research registries — Policy and academia
- Establish standardized registries recording diagnosis, sequence availability, scanner field strength, acquisition protocol, image quality, demographic variables, and clinical outcomes.
- The study demonstrates that missing sequences and scanner heterogeneity materially affect evaluation; these metadata should therefore be treated as first-class research variables.
- Dependencies: harmonized data standards, patient consent, privacy-preserving storage, sustainable funding, and governance involving local institutions.
- Training and education for clinicians and AI developers — Education
- Use the findings as a practical teaching case for dataset shift, class imbalance, external validation, shortcut learning, and the limitations of frozen foundation models.
- Graduate programs and hospital AI teams can reproduce the five-fold evaluation and examine how accuracy differs from MCC, macro-F1, and per-class recall.
- Dependencies: access to de-identified data and adequate technical supervision.
- Daily-life application: improved communication about AI-assisted MRI — Patients and the public
- The immediate public-facing implication is not a consumer diagnostic app, but clearer communication that an AI result from a model trained elsewhere may not generalize to local populations.
- Patients can be informed that AI outputs should complement, not replace, neurologist and radiologist assessment.
- Dependencies: responsible consent language, clinician communication, and avoidance of false reassurance or unnecessary alarm.
Long-Term Applications
The study’s findings support longer-term development of locally adapted, clinically validated systems rather than direct deployment of the evaluated frozen backbones.
- Parameter-efficient adaptation of foundation models to African MRI — Healthcare AI and software
- Apply LoRA, adapters, selective unfreezing, or prompt-based adaptation to Neuro-JEPA, BrainIAC, NeuroVFM, and Primus using Nigerian and broader African data.
- This would preserve much of the pretrained model while allowing adaptation to low-field scanners, local acquisition protocols, demographic variation, and population-specific disease presentations.
- Potential product: a site-calibration module that can be trained with a limited number of locally labeled scans.
- Dependencies: sufficient representative data, protection against overfitting, secure fine-tuning infrastructure, and external validation across hospitals and countries.
- African-inclusive neuroimaging foundation models — Research, industry, and global health
- Build new pretraining corpora containing MRI from multiple African regions, scanner manufacturers, field strengths, languages, clinical settings, and disease profiles.
- Pretraining should preserve geographic and clinical diversity rather than treating African data as a small post hoc calibration set.
- Potential product: a globally distributed neuroimaging backbone with documented performance by region and acquisition condition.
- Dependencies: federated or privacy-preserving data sharing, common imaging standards, sustainable funding, local ownership of data, and representation across countries and diagnostic groups.
- Federated multi-site learning for dementia and Parkinson’s disease — Healthcare systems and policy
- Train or adapt models across African hospitals without centralizing raw MRI, using federated learning or other privacy-preserving approaches.
- This could increase sample size while retaining institutional control over sensitive clinical data.
- Dependencies: reliable network and computing infrastructure, harmonized labels, secure aggregation, compatible preprocessing, and safeguards against site-specific shortcuts.
- Validated clinical decision-support systems — Healthcare
- After prospective studies, adapted models could support radiologists and neurologists by providing:
- disease-risk estimates;
- structured comparison with prior scans;
- sequence-specific evidence;
- uncertainty estimates;
- referral or follow-up prioritization.
- Systems should be designed as assistive tools with mandatory clinician review, especially because the study found low recall for some classes and possible shortcut learning.
- Dependencies: large prospective cohorts, calibrated uncertainty, clinically meaningful endpoints, regulatory approval, workflow integration, and evidence of improved patient outcomes rather than only higher benchmark scores.
- Multimodal fusion using T1w, T2w, and FLAIR — Neuroimaging software
- Develop models that explicitly handle missing modalities rather than relying on simple channel replication and zero masking.
- Possible approaches include modality-aware transformers, missing-modality imputation, sequence-specific encoders, and uncertainty-aware fusion.
- Potential tool: an inference system that produces a result when only one or two sequences are available while reporting how missing data affect confidence.
- Dependencies: paired multimodal scans, robust missing-data modeling, validation across acquisition protocols, and avoidance of imputation artifacts.
- Scanner and site harmonization — Radiology engineering and global health
- Create preprocessing and adaptation pipelines that normalize differences between 0.3 T, 1.5 T, and higher-field scanners, as well as differences in vendors, protocols, resolution, and noise.
- Such tools could include automated protocol checks, intensity harmonization, domain-adversarial training, and site-specific calibration.
- Dependencies: adequate samples from each scanner type, preservation of clinically relevant variation, and evidence that harmonization does not remove disease signals.
- Robustness testing against shortcut features — Responsible AI and regulation
- Establish standardized tests for borders, padding, skull-stripping errors, acquisition artifacts, and site identifiers.
- Models should be evaluated using counterfactual perturbations, artifact removal, region masking, and independent-site testing before clinical use.
- Potential regulatory artifact: a model card documenting performance and failure modes by modality, scanner, geography, and diagnostic class.
- Dependencies: access to raw acquisition metadata, expert review of attribution findings, agreed robustness standards, and regulatory acceptance of the evaluation framework.
- Population-specific calibration and fairness monitoring — Policy and healthcare governance
- Develop calibration methods and monitoring dashboards that track performance separately for age groups, sex, geography, scanner type, and diagnostic category.
- If a model performs well for dementia but poorly for Parkinson’s disease or controls, deployment thresholds should not be selected from aggregate accuracy alone.
- Dependencies: continuous post-deployment data, reliable labels, governance for updating models, and policies preventing performance disparities from being hidden by overall metrics.
- Large-scale longitudinal African neurodegenerative disease studies — Academia and public health
- Use standardized MRI and clinical follow-up to study disease progression, early conversion, treatment response, and regional variation in dementia and Parkinson’s disease.
- The resulting datasets could enable prognosis models rather than only three-way cross-sectional classification.
- Dependencies: long-term participant retention, consistent clinical assessments, funding for repeated imaging, and ethically governed data linkage.
- Resource-aware diagnostic systems for low-connectivity settings — Global health and daily life
- A future system could perform local inference on hospital hardware or portable workstations, synchronize only encrypted summaries, and provide decision support where specialist neurologists are scarce.
- This would be more practical than a cloud-only consumer application in regions with limited connectivity.
- Dependencies: validated compact models, reliable local computing, maintenance and cybersecurity, specialist referral pathways, and evidence that the system improves access without increasing misdiagnosis.
- Policy standards for equitable medical AI evaluation — Regulators and funders
- Require external validation on geographically and clinically diverse populations, reporting of per-class metrics, scanner-stratified results, missing-modality behavior, calibration, and shortcut analyses.
- Funding agencies and procurement bodies could make these criteria prerequisites for adopting imported neuroimaging foundation models.
- Dependencies: regulator capacity, shared reporting standards, access to representative data, and collaboration between African institutions, model developers, clinicians, and patient groups.
Glossary
- AUC-OVR: Area under the receiver-operating-characteristic curve computed using a one-versus-rest strategy. “The macro area under the receiver-operating characteristic curve (AUC-OVR) measures per-class separability independent of any decision threshold, computed in a one-versus-rest fashion.”
- Attribution map: A visualization indicating which input regions contributed to a model’s prediction. “We applied gradient-weighted class activation mapping (Grad-CAM) to obtain spatial attribution maps for model predictions.”
- Balanced accuracy: The average recall across classes, useful for imbalanced datasets. “Accuracy captures overall correctness, while balanced accuracy and the Matthews correlation coefficient (MCC) account for the moderate class imbalance in our cohort.”
- Backbone: The main feature-extraction component of a neural network to which another task-specific component is added. “We evaluate four pretrained 3D foundation models under a frozen-backbone protocol, alongside ViT3D as a from-scratch baseline.”
- BRISQUE: A no-reference image-quality metric that estimates visual distortion without requiring a pristine reference image. “We used BRISQUE (Blind/Referenceless Image Spatial Quality Evaluator) to identify and select the acquisition with the lowest BRISQUE score in the evaluation.”
- CLS token: A special transformer token whose representation summarizes an input sequence for classification or downstream tasks. “It produces patch-level representations without a CLS token and uses a sparse MoE architecture to encode multimodal brain MRI.”
- Contrastive learning: A self-supervised learning method that trains representations by bringing similar examples closer and separating dissimilar examples. “It is a ViT-B pretrained on 32,015 multiparametric brain MRIs from 16 datasets spanning 10 neurological conditions using SimCLR contrastive learning.”
- Cosine schedule: A learning-rate schedule in which the rate decreases according to a cosine-shaped curve. “The learning rate warms up linearly from 1\% to 100\% over five epochs, then decays following a cosine schedule to 1\% of the initial rate.”
- Cross-entropy: A classification loss that measures the difference between predicted class probabilities and the true labels. “The loss is cross-entropy with label smoothing of 0.1, which acts as a regularizer against overconfidence.”
- Domain adaptation: The process of modifying a model so that it performs effectively on data from a target domain that differs from its training domain. “Our results reinforce the view that frozen foundation models are not automatically sufficient in low-resource settings unless domain adaptation and dataset-specific considerations are taken into account.”
- Domain generalization: The ability of a model to perform well on domains or populations not represented in its training data. “\keywords{Brain MRI \and Foundation Models \and Domain Generalization \and Deep learning \and Explainability}”
- End-to-end training: Training all relevant model components jointly for the target task rather than keeping pretrained components fixed. “In contrast, the end-to-end trained ViT3D achieves higher accuracy and MCC on every task.”
- F1 score: The harmonic mean of precision and recall, measuring classification performance for a class. “Table~\ref{tab:per_class_f1} reports per-class F1 scores on the T1w+T2w configuration.”
- Field strength: The magnetic intensity of an MRI scanner, typically measured in tesla. “Data were acquired on scanners with field strengths ranging from 0.3 T to 1.5 T.”
- Fine-tuning: Further training of a pretrained model on a target dataset or task. “This is consistent with prior work showing that foundation models do not necessarily yield superior fine-tuned performance compared with standard supervised models in low-data settings.”
- FLAIR: Fluid-attenuated inversion recovery, an MRI sequence that suppresses fluid signals and can improve visualization of pathology. “Neuro-JEPA on FLAIR shows modest but still limited discrimination.”
- Foundation model: A large pretrained model intended to provide reusable representations for many downstream tasks. “Neuroimaging foundation models pretrained on large, predominantly western cohorts are increasingly proposed as general-purpose backbones for brain MRI analysis.”
- Frozen backbone: A pretrained feature extractor whose parameters are kept unchanged while another component is trained. “The pretrained weights are loaded from published checkpoints and remain frozen throughout training.”
- Grad-CAM: Gradient-weighted class activation mapping, a method for identifying image regions influential to a neural-network prediction. “We applied gradient-weighted class activation mapping (Grad-CAM) to obtain spatial attribution maps for model predictions.”
- Gradient-weighted class activation mapping: An interpretability technique that uses gradients to produce spatial importance maps for a predicted class. “For each subject, gradients of the predicted-class logit with respect to the model's spatial feature representations were used to estimate feature importance.”
- Label smoothing: A regularization technique that replaces hard class labels with slightly softened probability targets. “The loss is cross-entropy with label smoothing of 0.1, which acts as a regularizer against overconfidence.”
- Latent representation: An internal numerical encoding learned by a model to capture meaningful properties of input data. “It learns a shared latent representation across CT and MRI through masked volumetric prediction.”
- Linear probing: Training a simple linear classifier on fixed representations produced by a pretrained model. “For each pretrained backbone, we train a linear classifier on the frozen features.”
- LoRA: Low-Rank Adaptation, a parameter-efficient fine-tuning method that trains small low-rank updates instead of all model parameters. “We also did not evaluate lightweight adaptation strategies such as LoRA fine-tuning on the target population.”
- Macro-F1: The unweighted average of F1 scores computed separately for each class. “Macro-F1 averages per-class performance without weighting by support.”
- Matthews correlation coefficient: A classification metric based on all entries of the confusion matrix and robust to class imbalance. “Accuracy captures overall correctness, while balanced accuracy and the Matthews correlation coefficient (MCC) account for the moderate class imbalance in our cohort.”
- Mixture-of-experts: A neural architecture that routes inputs or tokens to selected specialist subnetworks rather than activating the entire model. “It is a ViT-B pretrained on 1,551,862 multimodal MRI scans (T1, T2, and FLAIR) using a JEPA objective, with mixture-of-experts routing.”
- Multimodal fusion: The combination of information from multiple imaging modalities within a model or analysis. “Future work should investigate parameter-efficient fine-tuning, shortcut-feature robustness, population-specific calibration, and richer multimodal fusion, particularly with T1+T2+FLAIR.”
- Patch embedding: The process of converting local image patches into vector representations suitable for a transformer. “For single-modality inputs the forward pass is straightforward: the volume goes through the backbone's patch embedding.”
- Patch token: A vector representing a local image patch after patch embedding. “Each model processes the input volume into a sequence of patch tokens that serve as input to the probing head.”
- Parameter-efficient fine-tuning: Adapting a pretrained model by updating only a small subset or compact representation of its parameters. “Future work should investigate parameter-efficient fine-tuning, shortcut-feature robustness, population-specific calibration, and richer multimodal fusion.”
- Pretraining: Initial training on a large dataset, often with a self-supervised objective, before adaptation to a target task. “Brain MRI foundation models are increasingly pretrained on large datasets drawn predominantly from well-resourced, non-African healthcare systems.”
- Probe: A lightweight task-specific model trained on representations generated by a fixed pretrained model. “This isolates the contribution of each model's learned representations from the capacity added by the probe.”
- Recall: The proportion of actual positive instances correctly identified by a classifier. “ViT3D achieves higher accuracy and MCC on every task (up to 53.4\% accuracy, MCC=0.27) and is the only model with non-trivial recall.”
- Self-supervised learning: Learning representations from data without manually supplied labels, typically through a constructed prediction task. “BrainIAC was introduced as a self-supervised foundation model for brain MRI designed to support broad downstream use.”
- Shortcut learning: Reliance on an easy but potentially spurious feature rather than clinically meaningful evidence. “A shortcut feature may support confident predictions without them being well grounded in pathology.”
- Skull stripping: Removing non-brain tissue from a brain image before analysis. “The dataset was skull-stripped using HD-BET.”
- Spatial attribution map: A spatial visualization of the image regions that influence a model’s output. “Subject-level attribution maps were averaged within each class and subsequently min--max normalized.”
- Transformer encoder: The attention-based component of a transformer that converts input tokens into contextualized representations. “The volume goes through the backbone's patch embedding, through the transformer encoder, and outputs a tensor of patch tokens.”
- Volumetric prediction: Prediction performed over three-dimensional medical-image volumes rather than two-dimensional images. “It learns a shared latent representation across CT and MRI through masked volumetric prediction.”
- Weight decay: A regularization method that penalizes large model parameters during optimization. “We optimize with AdamW at a learning rate of 0.001 and weight decay of .”
