- The paper introduces PathMoE, an interaction-aware mixture-of-experts model that fuses H&E whole-slide images, pathology reports, and nuclei-level cell graphs, achieving a macro-F1 of 0.799 on PBT cases.
- The model improves over image-only baselines by 0.037 macro-F1 on PBT and 0.041 on TCGA, while graph features provide useful structural knowledge when pathology text is unavailable or noisy.
- The gating mechanism identifies sample-specific modality contributions and helped correct four illustrated grade or lineage errors, although broader validation is needed to establish clinical interpretability and subtype robustness.
PathMoE addresses a persistent gap in computational pathology for pediatric central nervous system (CNS) tumors: while pathology foundation models have improved whole-slide image (WSI) analysis, most pipelines remain image-only and ignore the complementary evidence available in pathology reports and tissue microarchitecture. The paper proposes a multimodal framework that fuses H&E WSIs, report text, and nuclei-level cell graphs through an interaction-aware mixture-of-experts (I²MoE) architecture, delivering both measurable accuracy gains and sample-level interpretability (2603.01547).
Motivation and design rationale
Pediatric brain tumors are the leading cause of cancer-related mortality in children, and histopathological diagnosis is complicated by histologic heterogeneity, small training cohorts, and rare subtypes. Prior WSI methods—attention-based MIL, CLAM, TransMIL, state-space variants—extract strong morphological cues but struggle on difficult pediatric cases with subtle or overlapping morphology. Vision-LLMs such as CONCH, TITAN, PathChat, and SlideChat add textual context but are not tailored to the rare subtypes characteristic of pediatric neuro-oncology. PathMoE's premise is that explicit structural domain knowledge, in the form of nuclei cell graphs capturing spatial organization in the tumor microenvironment, can guide fusion of image and text modalities rather than serving as an isolated input.
Architecture
Each case yields three slide-level representations. Visual features come from UNIv2 patch embeddings at 20× magnification (256×256 tiles), aggregated via gated attention MIL. Textual features encode microscopic descriptions from pathology reports using TITAN by default. Cell graphs are built from 4× downscaled WSIs: nuclei are detected with HoverNet, encoded with ResNet34 features over 224×224 regions, connected into a k-NN graph (k=5), processed with GraphSAGE message passing, and pooled to graph-level features via attention MIL.
The core module applies I²MoE with five experts: three unimodal uniqueness experts, one synergy expert, and one redundancy expert. A perturbation strategy—replacing one modality with a random tensor across M forward passes—allows each expert's output to be decomposed so that the gating network can distinguish modality-specific information from shared or synergistic signals. A gating MLP over concatenated global representations produces softmax weights α, and final logits are the weighted sum of expert outputs. Training combines a classification loss with an interaction loss (λint=0.1) that regularizes expert specialization. Two fusion backbones are evaluated: Early Fusion (EF), a transformer over concatenated tokens, and SwitchGate (SG), an adaptive gating backbone derived from Switch Transformers.
Experimental setup
The primary evaluation uses the PBT dataset: 253 diagnostic H&E WSIs from 196 patients (age ≤ 29) collected under IRB approval, with 199 cases labeled into four classes (low-grade CNS tumor, high-grade CNS tumor, non-glial tumor, ependymoma). External validation uses 208 WSIs from 67 patients (age ≤ 29) drawn from TCGA-GBM and TCGA-LGG, covering astrocytoma, glioblastoma, and oligodendroglioma. All models use 10-fold cross-validation with patient-level splits to prevent leakage; macro-F1 is the primary metric. Image-only baselines (CLAM, TransMIL, S4MIL, MambaMIL, plus WSI-only EF and SG) share identical UNIv2 patch embeddings for controlled comparison.
Results
On PBT, PathMoE-EFWTG achieves the best macro-F1 of 0.799, versus 0.764 for the strongest image-only baseline (CLAM) and 0.762 for WSI-only EF—a gain of +0.037 from adding text and graph modalities. Class-level analysis shows where each modality contributes:
| Method |
Ependymoma F1 |
LG-CNS F1 |
HG-CNS F1 |
NGT F1 |
Macro F1 |
| CLAM |
.600 |
.780 |
.852 |
.824 |
.764 |
| EFW |
.600 |
.873 |
.787 |
.790 |
.762 |
| PM-EFWT |
.500 |
.881 |
.803 |
.868 |
.763 |
| PM-EFWTG |
.600 |
.914 |
.821 |
.862 |
.799 |
Adding report text substantially improves non-glial tumor recognition (F1 0.868 vs. 0.790), though partially at the cost of ependymoma performance—an honest trade-off the paper reports explicitly. Adding the graph modality on top of WT yields a further +0.036 to +0.039 macro-F1 boost across both backbones. Notably, SGWG improves ependymoma dramatically relative to SGW, and PathMoE-SGM0 lifts macro-F1 from 0.618 to 0.703 (+0.085), indicating text stabilizes representations when image-only gating is weak.
On TCGA, report text was unusable due to document noise, making this a test of graph knowledge alone. PathMoE-EFM1 reaches macro-F1 0.709 versus 0.668 for the best image-only model (+0.041), driven by large gains on the harder classes—oligodendroglioma F1 rises from 0.590 to 0.629 and astrocytoma from 0.586 to 0.665—while maintaining glioblastoma performance. This establishes structured microarchitectural priors as a critical signal precisely when text is unavailable or unreliable.
Interpretability analysis covers four cases where the image-only baseline confuses grade (HG ↔ LG) or lineage (HG ↔ NG); PathMoEM2 corrects all four, with non-trivial gate weights assigned to both graph (M3: 0.144–0.210) and text (M4: 0.167–0.184) experts. Independent review by a neuropathologist confirms that the corrected predictions align with spatial and architectural features consistent with what the graph expert encodes. The authors note these examples are illustrative rather than exhaustive.
A text encoder ablation shows domain-aligned pretraining outweighs scale: TITAN achieves 0.799 macro-F1 in the EFM5 setting, exceeding BioMistral (0.776)—a much larger 7B-parameter model—and CONCH (0.685). The same ordering holds under SG (0.742 vs. 0.714 and 0.666). Since TITAN builds on CONCH with additional WSI-oriented training, the result suggests histopathology-specific pretraining, not parameter count, drives fusion quality.
Limitations and open questions
Several constraints qualify these findings. The PBT cohort is small (199 classified cases across four classes), and per-class results—particularly ependymoma, which degrades in several multimodal configurations—are unstable, so the reported gains may not generalize uniformly across subtypes. On TCGA, only two of three modalities could be evaluated because report text was unusable, leaving the three-modality configuration untested externally. The interpretability claims rest on four qualitative cases plus gate-weight inspection; whether interaction weights constitute reliable clinical explanations at scale remains unverified. Finally, the framework depends on HoverNet segmentation quality and on foundation-model encoders whose biases propagate through fusion; the paper does not measure sensitivity to segmentation errors.
Conclusion
PathMoE shows that combining foundation-model visual and textual features with nuclei-graph structural knowledge through an interaction-aware MoE yields consistent macro-F1 improvements (+0.037 on PBT, +0.041 on TCGA) over strong image-only baselines, while its gating mechanism exposes per-sample modality contributions validated by neuropathologist review. The open questions concern robustness on rare subtypes, external validation of the full three-modality configuration, and the scalability of the interpretability evidence beyond case studies.