Duke Calcification Dataset v1 Benchmark
- The paper uses Duke Calcification Dataset v1 as an external test set to validate unsupervised domain adaptation, showing improved ROC AUC with CycleGAN-based methods.
- Duke Calcification Dataset v1 is a private mammography dataset of 788 biopsy-proven cases, featuring detailed lesion-level radiologist annotations and binary pathology labels.
- This benchmark highlights the challenges of vendor, imaging technique, and class-prior shifts, offering valuable insights for evaluating cross-domain generalization in calcification classification.
to=arxiv_search.search 娱乐开号ýyş av不卡免费播放 મુકદ્દمہ 人人中彩票 ,一本道 կազմվել json
{"query":"id:2607.06549 OR \"Unsupervised Domain Adaptation for Calcification Classification in Mammography Across Multi-Site Datasets\"","max_results":5,"sort_by":"submittedDate","sort_order":"descending"}
to=arxiv_search.search 北京赛车冠军 үрүм 盈立ынӡа json
{"query":"id:2402.09569 OR \"Automated Plaque Detection and Agatston Score Estimation on Non-Contrast CT Scans: A Multicenter Study\"","max_results":5,"sort_by":"submittedDate","sort_order":"descending"}
to=arxiv_search.search 】【。】【”】【json
{"query":"id:2211.01859 OR \"Computed tomography coronary angiogram images, annotations and associated data of normal and diseased arteries\"","max_results":5,"sort_by":"submittedDate","sort_order":"descending"}
to=arxiv_search.search 】【。】【”】【json
{"query":"id:2211.06728 OR \"Towards reliable calcification detection: calibration of uncertainty in coronary optical coherence tomography images\"","max_results":5,"sort_by":"submittedDate","sort_order":"descending"}
to=arxiv_search.search 】【。】【json
{"query":"id:2507.08214 OR \"Depth-Sequence Transformer (DST) for Segment-Specific ICA Calcification Mapping on Non-Contrast CT\"","max_results":5,"sort_by":"submittedDate","sort_order":"descending"}
Duke Calcification Dataset v1 is a private, in-house mammography dataset from Duke University Medical Center that was used as an independent external test cohort for malignant-versus-benign calcification classification in a multi-site domain adaptation study. In the documented arXiv usage, it comprises 788 biopsy-proven calcification cases acquired from 2008–2023, with lesion-level radiologist annotations and binary pathology labels, and it is evaluated only at test time rather than used for model training or validation [2607.06549].
1. Dataset identity and terminological scope
The explicit designation “Duke Calcification Dataset v1” appears in a mammography study on unsupervised domain adaptation for calcification classification across multi-site datasets, where it functions as the primary external U.S. test set [2607.06549]. In that study, the dataset is described as private, in-house, originating from Duke University Medical Center, and used to assess cross-site generalization under substantial variation in vendor, imaging technique, and examination type.
The name is significant because the broader calcification literature contains several datasets that might be confused with it but are explicitly not the same resource. A multicenter non-contrast chest CT plaque dataset for automated Agatston scoring is described without any Duke affiliation and without a formal standalone dataset name [2402.09569]. The Coronary Atlas / ASOCA training dataset is a CT coronary angiography resource from Australia and New Zealand and is explicitly stated not to be affiliated with Duke [2211.01859]. A non-contrast head CT cohort for segment-specific intracranial carotid artery calcification analysis is described only as a private clinical cohort and not as a Duke dataset [2507.08214]. An ex-vivo coronary OCT calcification detection dataset is reported from the University of Alabama at Birmingham rather than Duke [2211.06728]. This suggests that, within the cited arXiv corpus, the term refers specifically to a breast calcification mammography dataset, not to a CT or OCT vascular calcification dataset.
2. Cohort composition and reference labels
The dataset contains 788 biopsy-proven calcification cases. Of these, 247 cases (31.3%) are labeled malignant and 541 cases (68.7%) are labeled benign [2607.06549]. The label definition is pathology-based: a case is labeled malignant if malignancy is found on either the initial core needle biopsy or surgical excision pathology, when performed; otherwise it is labeled benign. The study also states explicitly that DCIS is classified as malignant.
All cases are calcification cases, and the study treats them at the lesion level. The dataset is not subdivided internally into train, validation, and test partitions for that paper’s experiments. Instead, all 788 cases form the external test cohort, while training and validation are performed exclusively on OPTIMAM [2607.06549]. The evaluation therefore probes out-of-domain generalization rather than within-site fitting.
The cohort is heterogeneous in acquisition context. The dataset spans 2008–2023 and includes a mixture of screening examinations (62.5%) and diagnostic examinations (37.5%). Within diagnostic exams, 26.9% of images are magnification views. This composition is described as having a relatively higher proportion of magnification and synthetic images than the public comparison datasets in the same study, making it a challenging external benchmark [2607.06549].
3. Imaging characteristics and annotation representation
The Duke cohort includes multiple vendors and image-generation techniques. Hologic accounts for 62.9% of images and GE for 37.1%. In terms of imaging technique, full-field digital mammography (FFDM) constitutes 58.6%, while synthetic 2D images derived from digital breast tomosynthesis constitute 41.4% [2607.06549]. This mixture is central to the dataset’s experimental role because it introduces both vendor shift and technique shift relative to the source-domain training data.
Lesion-level annotations were manually performed by a fellowship trained breast radiologist (11 years of experience) with radiology reports used as reference [2607.06549]. For the classification pipeline, each annotated calcification lesion is converted into one or more (512 \times 512) pixel patches. If a lesion is larger than 512 pixels in width or height, a sliding window strategy with overlapping windows is used to extract multiple patches.
At inference time, the model operates on these lesion patches rather than on entire mammograms. Patch-level malignancy scores are aggregated into a case-level score by maximum pooling:
$$
P_{\text{case}} = \max(P_{\text{patch}1}, P{\text{patch}2}, \ldots, P{\text{patch}_n})
$$
This case-level reduction rule is part of the study’s formal evaluation protocol on Duke Calcification Dataset v1 [2607.06549].
4. Role in the multi-site domain adaptation framework
Within the reported study, Duke Calcification Dataset v1 is used only for external validation / independent testing. It is not used in classifier training, in the training of the style-transfer models, or in hyperparameter tuning [2607.06549]. Instead, five models trained through 5-fold OPTIMAM cross-validation are applied to Duke, and the reported Duke metrics are averaged across those five models.
The framework evaluated on Duke has two components. The first is an unsupervised domain adaptation module based on AdaIN and CycleGAN. The second is a supervised classification module using Swin Transformer V2 as the backbone [2607.06549]. The style-transfer models are trained using OPTIMAM-derived data to generate vendor-specific and technique-specific training samples without additional annotations. Duke images themselves are not altered at inference time.
The rationale for using Duke as an external benchmark is the magnitude of the domain shift relative to OPTIMAM. OPTIMAM training lesions are described as predominantly Hologic FFDM screening calcification patches, whereas Duke includes 37.1% GE, 41.4% synthetic 2D, 37.5% diagnostic examinations, and 26.9% magnification views within diagnostic exams [2607.06549]. The label distribution also differs: OPTIMAM is reported as 73.0% malignant / 27.0% benign, while Duke is 31.3% malignant / 68.7% benign. This suggests that Duke was selected not merely as an external cohort, but as a deliberate stress test for vendor, technique, protocol, and class-prior shift.
5. Reported performance on the Duke cohort
Performance on Duke Calcification Dataset v1 is reported as case-level ROC AUC, using the maximum patch score per case. The study provides overall Duke performance and vendor-stratified performance for the baseline classifier, AdaIN-based domain adaptation, and CycleGAN-based domain adaptation [2607.06549].
| Setting | Duke subgroup | AUC |
|---|---|---|
| Baseline | Hologic-only | (0.67 \pm 0.02) |
| Baseline | GE-only | (0.70 \pm 0.03) |
| Baseline | All | (0.68 \pm 0.02) |
| AdaIN | Hologic-only | (0.69 \pm 0.01) |
| AdaIN | GE-only | (0.71 \pm 0.03) |
| AdaIN | All | (0.70 \pm 0.02) |
| CycleGAN | Hologic-only | (0.71 \pm 0.01) |
| CycleGAN | GE-only | (0.75 \pm 0.02) |
| CycleGAN | All | (0.73 \pm 0.01) |
The reported pattern is consistent across subgroups: both domain adaptation variants improve upon the baseline, and CycleGAN-based domain adaptation produces the highest AUC for Hologic-only, GE-only, and All Duke cases [2607.06549]. The overall Duke AUC increases from 0.68 for the baseline to 0.70 with AdaIN and 0.73 with CycleGAN.
The paper also states that the CycleGAN-based framework shows consistently higher sensitivity at corresponding specificity operating points than the baseline on Duke ROC curves, although exact sensitivity and specificity values are not tabulated numerically [2607.06549]. No formal statistical significance testing such as a DeLong test is reported for the Duke AUC differences.
6. Limitations, availability, and research significance
The study identifies several limitations that attach directly to the Duke dataset’s role as an evaluation resource. First, it is a private, single institution cohort [2607.06549]. While this provides a heterogeneous U.S. external test set relative to the U.K. training data, it does not by itself establish multi-institutional U.S. generalization. Second, although 788 cases is substantial for a curated lesion-level mammography dataset, the study notes broader concerns about bias from limited training and validation configurations, motivating its use of five-fold cross-validation on the source domain. Third, Duke includes a relatively high proportion of synthetic images and magnification views, but the paper does not report dedicated performance breakdowns by FFDM vs synthetic, screening vs diagnostic, or magnification vs non-magnification [2607.06549].
The dataset is described as private, in-house, and no public release mechanism, URL, or license is specified in the cited material [2607.06549]. Accordingly, its current role in the literature is closer to that of a controlled institutional benchmark than a public challenge dataset.
Its research significance lies in the combination of biopsy-proven labels, lesion-level radiologist annotation, and cross-domain heterogeneity. The paper explicitly positions it as a demanding external test environment for domain adaptation in calcification classification, with mixed GE/Hologic, FFDM/synthetic, and screening/diagnostic conditions [2607.06549]. A plausible implication is that Duke Calcification Dataset v1 is especially valuable for studying robustness to clinically realistic distribution shift rather than for optimizing within-site performance alone. In that sense, its main documented contribution is not merely its case count, but its use as a benchmark for evaluating whether calcification classifiers trained elsewhere retain discriminative performance in a substantially different clinical environment.