Hemorica: Unified Hemorrhage CT Dataset
- Hemorica is a publicly available non-contrast head CT dataset with unified annotations for five intracranial hemorrhage subtypes, enabling multi-task learning.
- It offers a hierarchical label structure including patient, slice, bounding box, 2D, and 3D segmentation annotations for precise lesion analysis.
- Baseline models demonstrate high classification and segmentation performance, highlighting its value for explainable AI and efficient volumetric reasoning.
Hemorica is a publicly available non-contrast head CT dataset for automated intracranial hemorrhage analysis that was introduced to support classification, detection, and segmentation within a single annotation framework. In its dataset paper, it is described as comprising 372 head CT examinations acquired between 2012 and 2024, with exhaustive annotations for five intracranial hemorrhage subtypes—epidural (EPH), subdural (SDH), subarachnoid (SAH), intraparenchymal (IPH), and intraventricular (IVH)—including patient-wise and slice-wise labels, subtype-specific bounding boxes, two-dimensional pixel masks, and three-dimensional voxel masks (Davoodi et al., 26 Sep 2025). Subsequent studies have used Hemorica both as a quantitative benchmark for class activation mapping and as a testbed for efficient slice-sequence detection pipelines, which situates it at the intersection of supervised segmentation, explainable AI, and low-compute volumetric reasoning (Rafati et al., 25 Aug 2025, Parvahan et al., 5 Jan 2026).
1. Scope and dataset identity
Hemorica was designed to address a recurrent limitation in intracranial hemorrhage research: public data are often fragmented across tasks, with one corpus supporting slice classification, another weak localization, and another full segmentation. The Hemorica dataset paper explicitly frames the resource as a unified, fine-grained benchmark for automated brain hemorrhage classification, segmentation, and detection (Davoodi et al., 26 Sep 2025).
The core acquisition description is specific. The scans are non-contrast head CT examinations acquired at Rasoul Akram Hospital in Tehran, Iran, exported in DICOM, converted to NIfTI, and anonymized. Axial slice counts per volume range from 16 to 56, with mean , and annotation was performed in the CT brain-window setting with HU and HU while preserving raw Hounsfield-unit values (Davoodi et al., 26 Sep 2025).
Hemorica is also defined by its subtype coverage. At the patient level, the dataset paper reports the following subtype breakdown: IPH , SAH , IVH , SDH , and EPH (Davoodi et al., 26 Sep 2025). This breadth distinguishes it from subtype-restricted resources such as HemSeg-200, which contains 222 non-contrast head CT volumes limited to IPH and IVH voxel annotation (Song et al., 2024). A plausible implication is that Hemorica is structured not only for binary hemorrhage detection but also for subtype-aware modeling and task cascades.
2. Annotation hierarchy and representational granularity
Hemorica’s principal technical characteristic is its multi-level label hierarchy. The dataset paper lists five annotation granularities: patient-wise classification labels indicating presence or absence of each subtype, slice-wise classification labels, subtype-specific bounding boxes derived from connected-component polygons, 2D pixel-level segmentation masks with one mask per subtype per slice, and 3D voxel-level segmentation volumes obtained by stacking 2D masks. It also provides per-subtype hemorrhage-volume estimates in voxels (Davoodi et al., 26 Sep 2025).
| Annotation level | Content |
|---|---|
| Patient-wise labels | Presence/absence of each subtype |
| Slice-wise labels | Slice-level hemorrhage labels |
| Detection labels | Subtype-specific bounding boxes |
| 2D segmentation | One mask per subtype per slice |
| 3D segmentation | Stacked voxel-level volumes |
This hierarchy is consequential methodologically. It supports binary or multi-label classification at coarse granularity, object-level detection via boxes, and dense lesion quantification via masks and volumes. The dataset paper explicitly recommends such use in multi-task learning, including unified models that perform slice-wise classification, bounding-box detection, and pixel segmentation in one pipeline, as well as curriculum learning that progresses from patient-level labels to slice-level labels, then bounding boxes, then 2D masks, and finally 3D volumes (Davoodi et al., 26 Sep 2025).
A common misconception is to treat Hemorica as a classification-only resource. That interpretation is contradicted by the published label inventory. Conversely, it is not merely a segmentation corpus: the presence of slice-level labels and subtype-specific boxes makes it suitable for weak supervision, explainability benchmarking, and detection-style formulations as well (Davoodi et al., 26 Sep 2025).
3. Annotation workflow and quality control
The dataset paper places strong emphasis on annotation governance. Two medically trained readers—a radiology specialist and a practitioner—annotated the scans independently while blinded to each other, and a neurosurgeon oversaw the protocol and adjudicated edge cases (Davoodi et al., 26 Sep 2025).
The workflow is described in four stages. First, 50 examinations were independently segmented by both readers in a pilot phase. Second, a neurosurgeon-led consensus phase established borderline rules and reduced inter-rater variability. Third, weekly annotation batches included both new volumes and a subset of prior cases for continuing intra-rater and inter-rater checks. Fourth, the radiologist’s labels were designated as ground truth, while the practitioner’s labels were retained for inter-rater agreement estimation (Davoodi et al., 26 Sep 2025).
The published quality statement is qualitative rather than coefficient-based. The paper reports “low” inter-rater variability, confirmed by repeated sampling and neurosurgeon review, but does not state an exact value (Davoodi et al., 26 Sep 2025). This is significant because many medical-imaging resources provide dense annotations without a detailed adjudication mechanism. Here, the double-reading workflow, the pilot consensus phase, and the neurosurgeon’s role are presented as the principal controls on label stability.
This annotation design has downstream implications for benchmark credibility. Because the dataset supports both lesion contours and derived detection labels, variability at the mask level propagates into box-level and volume-level tasks. The paper’s emphasis on repeated checking suggests an effort to stabilize all three derived representations simultaneously, not only the final binary mask (Davoodi et al., 26 Sep 2025).
4. Statistical properties and clinical realism
The dataset paper presents Hemorica as clinically realistic in both case composition and lesion characteristics. At the patient level, Hemorica contains 109 healthy cases and 263 hemorrhagic cases 0. The paper compares these proportions with RSNA 1 healthy, 2 hemorrhagic3 and CQ500 4 (Davoodi et al., 26 Sep 2025).
At the slice level, it reports 9,388 healthy slices 5 and 2,678 hemorrhagic slices 6, with subtype distribution IPH 7, SAH 8, IVH 9, SDH 0, and EPH 1 (Davoodi et al., 26 Sep 2025). Lesion-area statistics are also provided: IPH masks have mean size 3,076 voxels, median 1,985, and maximum 27,322, whereas IVH and SAH have substantially smaller mean and median areas, reflecting segmentation difficulty (Davoodi et al., 26 Sep 2025).
The intensity analysis is equally relevant. Hemorrhagic voxels are described as having a unimodal peak at high HU, well separated from the normal brain distribution spanning 2 to 3 HU. The paper contrasts this with PhysioNet-ICH, where a second lower-HU peak is observed, suggesting subacute-versus-acute bleed differences (Davoodi et al., 26 Sep 2025). This matters because HU-distribution shape influences both threshold-based preprocessing and representation learning.
A notable issue is that later Hemorica-based studies do not report numerically identical corpus sizes. The CAM benchmarking paper describes Hemorica as 327 non-contrast head CT studies with 12,067 axial slices, of which 2,679 are hemorrhage-positive and 9,388 hemorrhage-negative (Rafati et al., 25 Aug 2025). The dataset paper, by contrast, reports 372 examinations and 9,388 healthy versus 2,678 hemorrhagic slices (Davoodi et al., 26 Sep 2025). This suggests a versioning difference or differing inclusion criteria, but no harmonizing explanation is stated in the supplied descriptions.
5. Baseline models and reference performance
The dataset paper establishes baseline results for two core tasks: binary slice classification and binary lesion segmentation. For classification, all models are ImageNet-1K pretrained. The architecture set includes ResNet-18/50, DenseNet-121/161, EfficientNetV2-Small/Large, Swin Transformer V2 Tiny/Small, and MobileViT-XS/S. Input preprocessing is fixed: a single slice is clipped to 4, normalized, replicated to three channels, and resized to 5. Training uses a patient-wise 6 split preserving subtype ratios, cross-entropy loss, Adam, batch size 16, 50 epochs, and learning rates 7 for CNNs and 8 for ViTs, with threshold 9 at inference (Davoodi et al., 26 Sep 2025).
The headline classification baseline is MobileViT-XS with precision 0, recall 1, 2, AUC 3, and specificity 4. Across all ten classification models, the mean 5 is approximately 6 and AUC exceeds 7 (Davoodi et al., 26 Sep 2025).
For segmentation, the paper evaluates U-Net and PSPNet decoders paired with the same pretrained encoders, using identical windowing and preprocessing, pixel-wise cross-entropy, Adam, batch size 16, 50 epochs, and the same train/validation split. The best reported model is U-Net with DenseNet161 encoder, achieving precision 8, recall 9, Dice 0, and IoU 1; mean Dice across all U-Net variants is approximately 2 (Davoodi et al., 26 Sep 2025).
These baselines position Hemorica differently from weakly supervised NCCT work. For example, a ResNet-LSTM plus CAM, K-means pseudo-masks, and 3D U-Net pipeline trained only from image-level slice labels reported Dice 3 on INSTANCE validation data, despite low contrast and poor SNR in NCCT (Ramananda et al., 2023). Hemorica’s strong fully supervised baselines therefore provide an upper-anchor for studies that seek to recover localization or segmentation from weaker labels.
6. Explainability, efficient detection, and research uses
Hemorica has already been used for benchmarking beyond conventional supervised segmentation. One explainability study evaluates nine CAM methods—GradCAM, HiResCAM, GradCAM-ElementWise, GradCAM++, XGradCAM, AblationCAM, EigenCAM, EigenGradCAM, and LayerCAM—across multiple EfficientNetV2-S stages. In that benchmark, the strongest localization occurs at stage 5, corresponding to layer 4; HiResCAM achieves the highest bounding-box alignment with BBox Dice 5 and BBox IoU 6, while AblationCAM attains the best pixel-level overlap with Pixel Dice 7 and Pixel IoU 8 (Rafati et al., 25 Aug 2025). The benchmark is technically important because the models were trained solely for classification rather than segmentation supervision.
A second direction treats Hemorica as a detection corpus for lightweight volumetric reasoning. A 2026 study reformulates CT volumes as sequential video streams along the axial direction and combines a YOLO Nano detector with ByteTrack to enforce slice-to-slice consistency. In that work, YOLOv11n is selected as the backbone with 9, and the final Hybrid ByteTrack pipeline raises detection precision from 0.703 to 0.779 while maintaining recall, reaching 0 on the independent test set (Parvahan et al., 5 Jan 2026). This use case illustrates that Hemorica’s box and slice labels support not only dense segmentation but also edge-oriented detection pipelines.
The dataset paper itself recommends several downstream uses: multi-task learning, curriculum learning, transfer learning from Hemorica to larger weakly labelled cohorts such as RSNA and Qure25k, and AI-assistant design using segmentation masks for automated volume quantification and decision-support systems that flag slices as “high-risk hot spots” while providing subtype and volume estimates (Davoodi et al., 26 Sep 2025). Taken together, these uses indicate that Hemorica is best understood as an infrastructure dataset for hierarchical supervision rather than a single-task benchmark.
The main caution in reading the current literature is that published Hemorica descriptions are not yet numerically uniform across papers, particularly with respect to examination and slice counts (Rafati et al., 25 Aug 2025, Davoodi et al., 26 Sep 2025). Even so, the corpus’s defining contribution is stable across descriptions: it provides a rare combination of slice labels, subtype-specific detection targets, and high-quality segmentation masks within a single non-contrast head CT resource, enabling rigorous evaluation of classification, localization, segmentation, explainability, and computational-efficiency trade-offs.