Papers
Topics
Authors
Recent
Search
2000 character limit reached

Automated Histopathology Report Generation via Pyramidal Feature Extraction and the UNI Foundation Model

Published 18 Feb 2026 in eess.IV, cs.AI, and cs.CV | (2602.16422v1)

Abstract: Generating diagnostic text from histopathology whole slide images (WSIs) is challenging due to the gigapixel scale of the input and the requirement for precise, domain specific language. We propose a hierarchical vision language framework that combines a frozen pathology foundation model with a Transformer decoder for report generation. To make WSI processing tractable, we perform multi resolution pyramidal patch selection (downsampling factors 23 to 26) and remove background and artifacts using Laplacian variance and HSV based criteria. Patch features are extracted with the UNI Vision Transformer and projected to a 6 layer Transformer decoder that generates diagnostic text via cross attention. To better represent biomedical terminology, we tokenize the output using BioGPT. Finally, we add a retrieval based verification step that compares generated reports with a reference corpus using Sentence BERT embeddings; if a high similarity match is found, the generated report is replaced with the retrieved ground truth reference to improve reliability.

Summary

  • The paper presents a modular report-generation system that combines pyramidal patch selection, quality filtering, frozen UNI embeddings, and a six-layer Transformer decoder instead of end-to-end MLLM fine-tuning.
  • The method achieved a composite score of 0.8093, ranking eighth of 24 REG 2025 teams and within 4.7% of the winner, while reducing training GPU memory from about 16 GB to 4 GB through feature caching.
  • The system produces consistent diagnostic templates and suppresses hallucinations with retrieval verification, but still struggles with invasive-versus-in-situ distinctions, prostate grading, rare diagnoses, and validation beyond one benchmark.

Overview

This paper presents a modular vision–language framework for Automated Histopathology Report Generation (AHRG), developed by the MedInsight-ViseurAI team for the REG 2025 Grand Challenge. Rather than fine-tuning an end-to-end Multimodal LLM (MLLM), the system couples a frozen pathology foundation model encoder (UNI, ViT-Large/16) with a compact 6-layer Transformer decoder, and adds retrieval-based verification to suppress hallucinated outputs. The method achieved a composite ranking score of 0.8093 in Test Phase 2, ranking 8th of 24 teams and within 4.7% of the winner (2602.16422).

Method

The pipeline comprises three stages: pyramidal patch selection with quality filtering, feature extraction via UNI, and text decoding with retrieval post-processing.

Pyramidal patch selection. The framework processes WSI pyramid levels ℓ∈{6,5,4,3}\ell \in \{6,5,4,3\} in a coarse-to-fine manner. At each level, tissue masks are computed by HSV thresholding (τS=20\tau_S = 20, τV=30\tau_V = 30) followed by morphological opening/closing with a 5×55\times5 structuring element. Non-overlapping 256×256256 \times 256 candidate patches are retained when tissue coverage exceeds 10%.

Quality filtering. Three interpretable criteria remove non-informative patches: Laplacian-variance focus scoring (reject if below 40), HSV Value/Saturation exposure checks (μV∉[40,245]\mu_V \notin [40,245] or μS<12\mu_S < 12), and dark-pixel fraction (>0.2>0.2 rejected as dust or pen-mark artifacts). A stratified random sampling scheme caps each slide at Nmax=2500N_{max} = 2500 patches, allocated proportionally across pyramid levels.

Frozen foundation-model encoding. Each selected patch is encoded by UNI—a ViT-L/16 distilled via DINOv2 on over 100 million histopathology patches—yielding a 1024-dimensional CLS-token embedding per patch. Keeping all 307M encoder parameters frozen reduces GPU memory from roughly 16 GB to 4 GB during decoder training, and permits pre-computation and caching of features in HDF5 format, fully decoupling extraction from training.

Decoder. A lightweight projection layer maps patch features into a 1024-dimensional decoder space, serving as cross-attention memory with key-padding masks for variable-length inputs. The 6-layer decoder uses 8 attention heads, feed-forward dimension 2048, dropout 0.1, sinusoidal positional encodings, and a maximum sequence length of 64 tokens. Text is tokenized with the BioGPT vocabulary (42,384 tokens), reducing subword fragmentation of biomedical terminology. Training minimizes teacher-forced cross-entropy using AdamW with a two-phase schedule: a 10-epoch warmup at 5×10−55\times10^{-5} decaying to a base rate of τS=20\tau_S = 200, over 350 epochs at batch size 64.

Retrieval-based verification. Generated reports are embedded with Sentence-BERT (all-MiniLM-L6-v2, 384 dimensions). If cosine similarity to the nearest ground-truth report in the training corpus exceeds τS=20\tau_S = 201, the generation is replaced by the retrieved reference; otherwise the original generation is kept.

Evaluation

Experiments use the REG 2025 Grand Challenge dataset: 10,494 WSI–report pairs from five institutions spanning seven organ systems, split into 8,494 training samples and two 1,000-sample test sets with strict patient-level separation. The challenge's composite score weights keyword Jaccard similarity most heavily (0.4), semantic embedding similarity next (0.3), and ROUGE/BLEU combined only 0.15, reflecting clinical prioritization of diagnostic terminology over stylistic overlap.

Rank Team Score
1 IMAGINE Lab 0.8494
8 MedInsight-ViseurAI (ours) 0.8093
9 ADCT 0.8040

Qualitative analysis shows exact or near-exact matches for common entities—invasive breast carcinoma NST grade II, colonic chronic inflammation, lung squamous cell carcinoma—and consistent adherence to the canonical [Organ], [biopsy type]; [diagnosis] template. Two failure modes recur: confusion of invasive versus in situ breast lesions with multi-attribute descriptors, and misgrading of prostate adenocarcinoma (predicted Gleason 6 (3+3) instead of the correct 7 (3+4)). Minor procedure-name discrepancies (e.g., "colposcopic" vs. "punch" biopsy) occur but do not alter diagnostic conclusions.

Discussion

The authors argue that architectural simplicity plus careful training procedure can compensate substantially for reduced capacity relative to MLLM approaches. Three claims stand out:

  • Structural consistency: because generation follows learned templates deterministically rather than sampling freely, the authors report virtually no instances of format violations or out-of-domain text—an advantage for clinical deployment where standardized structures are mandatory.
  • Efficiency: freezing the encoder and caching features enables iterative experimentation within resource-constrained settings, avoiding the billions of parameters typical of end-to-end MLLM fine-tuning.
  • Domain tokenization: BioGPT vocabulary shortens effective sequences for pathological terms and strengthens visual-to-diagnostic-phrase associations.

The strong performance on Gleason-style grading systems is attributed to their prevalence in training data and consistent linguistic templates; errors concentrate where multiple semi-independent attributes create combinatorially sparse supervision.

Limitations and open questions

The paper concedes several constraints directly. Ground-truth labels were unavailable for the full test set, so no quantitative per-category error analysis was possible—the qualitative observations rest on spot checks rather than systematic measurement. The retrieval-correction mechanism assumes that a similarity above τS=20\tau_S = 202 implies a reliable reference exists; this assumption may suppress valid rare diagnoses underrepresented in the training corpus, effectively biasing output toward common cases. Evaluation is confined to a single challenge dataset drawn from specific institutional contexts, leaving generalizability untested. Finally, the framework generates only diagnostic summary components; gross descriptions and ancillary-test recommendations are out of scope. Open questions include whether structured prediction heads or auxiliary attribute-level objectives would resolve multi-attribute grading failures, and whether the reported robustness transfers beyond the REG 2025 distribution.

Conclusion

The paper demonstrates that competitive automated histopathology report generation is achievable with a frozen pathology foundation model, hierarchical patch selection, domain-specific tokenization, and lightweight decoder training, reaching a top-ten result on REG 2025 without MLLM-scale compute. Its main residual weaknesses are fine-grained multi-attribute grading, potential suppression of rare diagnoses through retrieval replacement, and single-benchmark validation.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.