- The paper introduces MixTIME, a router-gated mixture-of-experts model that combines image, text, and transcriptomic pathology foundation models to predict 17 multiplex immunofluorescence proteins from H&E whole-slide images at pixel-level resolution.
- MixTIME achieves state-of-the-art biomarker prediction across HEMIT and ORION, while its predicted protein maps improve spatial clustering and survival modeling compared with image-only and transcriptomic baselines.
- The paper shows that imputed biomarkers can support pathology reports and longitudinal tumor microenvironment analysis, but low-expression markers, limited clinical validation, and reduced report conciseness remain important limitations.
MixTIME is a multimodal mixture-of-experts (MoE) foundation model that predicts multiplex immunofluorescence (mIF) protein expression directly from hematoxylin and eosin (H&E) whole-slide images (WSIs) at pixel-level resolution. Rather than training a single image-to-protein regressor, the model composes frozen pathology foundation models (PFMs) trained on distinct modality pairs—image-only (UNIv2), image-text (CONCHv1.5), and image-transcriptomic (STPath)—and injects their knowledge into a fine-tuned mIF predictor backbone (GigaTIME) via cross-attention, with a learnable router weighting each expert's contribution. Trained with a distribution- and tendency-aware loss combining channel-level and pixel-level negative Pearson correlation with Smooth L1, MixTIME achieves state-of-the-art mIF prediction across 17 protein markers on two benchmark datasets and demonstrates that the predicted biomarkers transfer to spatial domain identification, survival prediction, pathologist-validated report generation, and longitudinal tumor microenvironment (TIME) analysis (2606.18123).
Architecture and training objective
The model treats a WSI as a collection of patches and defines the task as predicting a per-patch mIF tensor m^k∈RC×H×W, where C is the number of protein channels. Four experts participate: GigaTIME serves as the encoder-decoder backbone providing spatial feature maps; the three frozen PFMs (UNIv2, CONCHv1.5, STPath) inject complementary representations through criss-cross attention modules. A learnable router modulates expert contributions, and only the backbone plus fusion modules are fine-tuned. The loss combines a channel-wise Pearson term (consistency across protein channels at each spatial location), a pixel-wise Pearson term (correct spatial distribution per channel), and Smooth L1 magnitude loss, with λch=λpx=1.0. Training uses AdamW (learning rate 1×10−4, weight decay 1×10−5, batch size 16, up to 100 epochs) on a single NVIDIA H200.
The authors evaluate on HEMIT (3 markers, smaller scale) and ORION (17 markers, larger scale), using pixel-level Pearson (PCC) and Spearman (SCC) correlation coefficients against GigaTIME (zero-shot, layer-finetuned, and fully fine-tuned variants), HEX, MIPHEI-ViT, and ROSIE. MixTIME ranks first in aggregate PCC across both datasets and most biomarkers, with two notable caveats stated plainly by the authors: it does not outperform baselines on the background channel Hoechst, and for intrinsically difficult, low-expression proteins no method—including MixTIME—achieves substantial improvement, which the authors attribute to mIF measurement quality (staining, signal sparsity, batch effects) rather than model capacity. Because MixTIME is natively pixel-level, it also avoids the patch-to-pixel approximation steps required by HEX and ROSIE, yielding higher effective resolution. Case-study visualizations on both datasets show predicted maps closely tracking measured expression for well-predicted proteins.
Ablations on HEMIT support the design choices: using all four experts outperforms any reduced expert set on both PCC and SCC; feeding STPath's full gene expression profile (rather than a low-dimensional embedding) into the router is superior; and the mixed PCC + Smooth L1 loss outperforms alternatives that capture only magnitude or only correlation-level similarity, consistent with prior findings in multimodal biomarker prediction (2606.18123).
Downstream gains in clustering and survival prediction
Because mIF is expensive and rarely available, the practical value of MixTIME lies in whether its predictions act as a useful surrogate modality. For spatial domain identification on four human-annotated datasets from STImage-1k4M, MixTIME-derived embeddings (enriched with image, text, transcriptomic, and predicted mIF information) outperform UNIv2, GigaPath, Triplex, BLEEP, and STPath across all four clustering metrics (ARI, AMI, homogeneity, NMI), with variance comparable to baselines and consistent gains across Leiden resolutions from 0.2 to 2.0. UMAP case studies show better separation of cortical layer 6 and white matter spots than image-only baselines. For weakly supervised survival prediction on four TCIA cohorts (MBC, SURGEN, HNSC, LUAD) under a multiple-instance learning framework, adding MixTIME biomarkers to PFM features yields consistent C-index improvements, most pronounced in HNSC. These results indicate that predicted protein expression carries information not already captured by image or transcriptomic embeddings alone.
Pathologist-validated biomarker-enhanced report generation
The authors integrate MixTIME-predicted biomarker ranks (normalized, sorted by expression level) into GPT-5-based pathology report generation and evaluate blinded, with pathologists from multiple institutions scoring 11 reports across Completeness, Relevance, Conciseness, Coherence, and Clarity. The MixTIME-enhanced reports perform competitively or strongly, particularly on Relevance and Completeness, but the authors are explicit that the enhancement does not uniformly dominate: it is more variable on Clarity and measurably reduces Conciseness (statistically significant by Wilcoxon rank-sum), because additional molecular context lengthens reports. Two further findings temper enthusiasm: when inter-pathologist variability is ignored, mIF integration only matches other methods on Completeness; and GPT-5 direct generation did not score significantly below original human annotations, validating the base generator as much as the biomarker augmentation. Pathologists also diverge in acceptance—one endorsed biomarker inclusion, another objected that histology-plus-rank conclusions cannot be drawn without exact expression levels—suggesting biomarker-enriched reports suit molecularly oriented readers, and that reporting precise intensities rather than ranks is an open design question. The experiment is preliminary in scale (10 ROIs from one WSI, a small evaluator panel), a limitation the authors acknowledge.
Biomarker discovery and longitudinal TIME profiling
Two translational analyses exploit MixTIME's ability to impute proteins where mIF was never measured. First, on a public cutaneous squamous cell carcinoma spatial transcriptomics cohort with peri-neural invasion (PNI) and anti-PD-1 resistance annotations, the authors correlate predicted mIF intensity for 16 proteins against three gene signature programs (Cancer-Induced Nerve Injury, Anti-Tumoral Immunity, Immunosuppression). The strongest signals are structural rather than single-cell co-expression: FOXP3 (Treg marker) correlates with PECAM1 (endothelial), consistent with perivascular Treg niches; FOXP3 correlates with IRF8 (APC activation), interpreted as coexisting activation and suppression; and CD45RA (naive T cells) correlates positively with CD163 (M2-like TAMs), a pattern the authors link to immune suppression in resistant tumors. These are hypothesis-generating correlations from predicted—not measured—protein data, and the authors do not validate them against orthogonal protein assays in this cohort.
Second, on WSIs from Harvard Medical School patients with paired EHR timelines (one lung adenocarcinoma patient at three time points; a sarcomatoid mesothelioma patient at four), MixTIME tracks protein dynamics across clinical events: elevated epithelial/proliferative and immune markers (Pan-CK, E-cadherin, Ki67, PD-L1, FOXP3) at diagnostic biopsy, convergence toward baseline after lobectomy with tumor-negative nodes, and a distinct immune/stromal remodeling pattern (SMA, CD68, CD20, CD4, CD3e, CD163) at a later new nodule. The mesothelioma case similarly distinguishes an extensively invasive resection (broadly elevated CD45, CD68, CD20, CD4/CD8a, CD31) from a more contained local recurrence. The predicted trajectories align with the EHR narrative, supporting the use of imputed mIF as a severity and disease-state readout, though this constitutes a small number of single-patient case studies rather than a validated clinical assay.
Limitations and open questions
The paper concedes three principal limitations. Prediction accuracy remains limited for low-expression or noisy biomarkers, reflecting both biology and technical variation in mIF measurement (staining quality, batch effects, tissue heterogeneity). Training and evaluation rely on public datasets (HEMIT, ORION); validation across additional cancer types, institutions, scanners, staining protocols, and populations is required before clinical deployment. The report-generation study is small in both ROIs and evaluators, and the finding that biomarker context reduces conciseness implies that deployment should be conditional rather than default. Open questions include whether reporting exact mIF intensities (rather than ranks) improves pathologist acceptance, whether the FOXP3/PECAM1 and CD45RA/CD163 associations hold against measured protein data, and how router weights generalize across tissue types and staining protocols.
Conclusion
MixTIME demonstrates that a router-gated composition of frozen PFMs spanning image, text, and transcriptomic modalities improves pixel-level mIF prediction from H&E over single-modality baselines, and that the resulting predicted protein maps function as a usable surrogate modality for clustering, survival modeling, report generation, and longitudinal TIME analysis. The strongest evidence is the consistent benchmark superiority on mIF prediction and clustering; the translational analyses are suggestive but rest on predicted rather than measured proteins and small clinical cohorts.