Cross-context generalization of virtual staining models

Determine whether the performance and model-family rankings of unsupervised H&E-to-Sirius Red virtual-staining architectures generalize to other tissue preparations, disease contexts, and human NASH or NAFLD liver biopsies.

Background

The study evaluates six unsupervised image-to-image translation architectures using a single bile-duct ligation mouse-liver cohort processed in one staining laboratory. Consequently, the reported scaling behavior, task-specific CPA accuracy, perceptual quality, and ensemble-based uncertainty may depend on the tissue preparation, disease model, and acquisition conditions used in that cohort.

The authors explicitly note that generalization to other tissue preparations and disease contexts has not yet been established, and that fibrosis patterns, collagen architecture, and staining behavior in human NASH or NAFLD biopsies may alter the transferability of the observed architectural rankings. Independent cross-cohort validation is therefore required.

References

The dataset covers a single tissue type, disease model, and staining laboratory; while the WSI scanning pipeline is well-standardised (reducing one common source of cross-centre variability), generalisation to other tissue preparations and disease contexts remains to be validated. The bile-duct ligation cohort is a preclinical mouse model, and fibrosis patterns, collagen architecture, and staining behaviour in human NASH or NAFLD biopsies may differ in ways that affect ranking transfer.

Towards Reliable AI-Based Histological Staining: A Systematic Study of Scaling and Uncertainty in Unpaired Generative Models  (2608.24626 - Siddiqui et al., 25 Aug 2026) in Section 9, Discussion, subsection “Limitations”