---
title: 'OPTIMAM: Mammography Research Database'
url: https://www.emergentmind.com/topics/optimam
type: topic
---

# OPTIMAM: Mammography Research Database

OPTIMAM is a Cancer Research UK–funded mammography research programme, originally created to support **OPTIMAM (2008–2013)** and **OPTIMAM2 (2013–2018)**, whose central data asset is the **OPTIMAM Mammography Image Database (OMI-DB)**: a large, curated, shareable research database of mammography images linked to screening, clinical, and pathological information [2004.04742]. OMI-DB consists of several relational databases and cloud storage systems and, in the main content description, contains **2,889,312 images** of all types from **173,319 women**; the abstract summarizes this as **over 2.5 million images** from the same cohort [2004.04742]. Its combination of mammographic imaging, follow-up, interval cancers, prior examinations, and lesion-level expert marks has made OPTIMAM a recurrent source domain for computer-aided detection, risk prediction, image perception studies, virtual clinical trials, and cross-site generalization research [1902.07323].

## 1. Programme origins and population scale

OPTIMAM was built to address a structural limitation in medical-imaging research: the lack of very large, well-curated, sharable datasets with trustworthy outcome labels and longitudinal follow-up [2004.04742]. Within the programme, OMI-DB was designed not only for the original OPTIMAM studies but also for broader use in **virtual clinical trials (VCT)**, **computer-aided detection (CAD)**, **artificial intelligence / machine learning**, **image perception studies**, **reader training**, and **quality assurance** [2004.04742]. The database was collected from **three UK breast screening centres**: **Jarvis Breast Screening Centre, Guildford**; **St George’s Hospital, South West London**; and **Addenbrooke’s Hospital, Cambridge** [2004.04742].

The cohort reported in the abstract is stratified as follows [2004.04742]:

| Group | Women |
|---|---:|
| Normal breasts | 154,832 |
| Benign findings | 6,909 |
| Screen-detected cancers | 9,690 |
| Interval cancers | 1,888 |

The collection strategy combined exhaustive ascertainment and sampling. It included **all screen-detected cancers** from the three sites, **prior screening mammograms of interval cancers since 2012**, **all women screened between 1 January 2014 and 31 December 2014**, and a **random 25% sample of all women screened in 2012, 2013 and 2015** [2004.04742]. Collection is **ongoing**, and women are **followed up** with their clinical status updated according to subsequent screening episodes [2004.04742]. This longitudinal design makes OPTIMAM more than a cancer archive; it is a screening-population resource containing normal, benign, screen-detected, and interval-cancer trajectories.

## 2. Database structure, imaging content, and lesion annotation

OMI-DB stores both images and linked metadata extracted largely from the **National Breast Screening System (NBSS)**, including **radiological information**, **clinical information**, and **pathological information** [2004.04742]. The imaging component includes **screening mammograms**, both **processed** and **unprocessed** images, initially **2D digital mammography**, and later additional modalities such as **tomosynthesis** and **MRI** [2004.04742]. When images are loaded, relevant DICOM tags are extracted into a searchable index; fields shown in the schema include `study_id`, `series_id`, `modality`, `manufacturer`, `view_position`, `image_laterality`, and `presentation_intent_type` [2004.04742].

The relational schema links the hierarchy from woman to image. The simplified schema includes entities such as **client**, **episode**, **episode_event**, **imaging**, **study / series / images**, **lesion**, **marks**, and **clinical/surgery/biopsy data** [2004.04742]. Representative fields shown in the paper include `ClientID`, `EpisodeID`, `EventID`, `LesionID`, `Episode Type`, `EpisodeAction`, `EpisodeOpenedDate`, `EpisodeClosedDate`, `AssessmentType`, `DatePerformed`, `FinalAction`, and `MammoOpinion`, together with lesion-side, lesion-position, and pathology-related fields such as **in-situ grade** and **invasive grade** [2004.04742]. This structure supports linkage across screening history, biopsy results, surgery, and image-level annotations.

A major value-added feature is lesion localization. Because lesion location is **not routinely stored in clinical databases**, it was collected specifically for OMI-DB by experienced mammography readers, who manually annotated **lesion location**, **lesion area**, **radiological appearance**, and **conspicuity** [2004.04742]. The paper reports that **7,143 lesions have been marked**, including **approximately 60% of screen-detected cancers** [2004.04742]. The schema’s `marks` table includes fields such as `mark_id`, `ImageID`, `mark_type`, `mark_shape`, `radiological_appearance`, and `conspicuity` [2004.04742]. This makes OPTIMAM suitable not only for whole-exam classification but also for lesion localization and detection studies.

## 3. Ingestion workflow, de-identification, and controlled access

OMI-DB is architected as **several relational databases and cloud storage systems**, with distributed collection and centralized storage for research use [2004.04742]. Two collection modes are described: **automated remote-site collection** and **stand-alone collection** [2004.04742]. In the automated mode, a physical or virtual server on site queries NBSS for women matching collection requirements such as date range, outcome classification, or high-risk status, and then retrieves images and clinical data for screening and assessment episodes [2004.04742]. In the stand-alone mode, a tool can ingest images from a prepared folder while querying NBSS for the relevant clients [2004.04742].

At collection time, imaging and screening data are **pseudonymised**, lookup tables are created at the collection site, and images, image metadata, and screening data are uploaded to the cloud for storage in the central database [2004.04742]. For external sharing, the resource applies **additional de-identification**, and a dedicated database records which cases were shared with each third party, together with investigator and sharing metadata [2004.04742]. All collection activity is logged, including NBSS queries, PACS queries, metadata on collection runs, and client-related collection activity [2004.04742]. A web-enabled application, **MedxViewer**, supports remote case viewing, feature annotation, and observer studies [2004.04742].

Governance is integral to the OPTIMAM model. The project has approval from an **ethical research committee** specialising in research databases, organised by the **NHS Health Research Agency** [2004.04742]. **Cancer Research UK (CRUK)** retains the **intellectual property** of the database, and sharing occurs through agreements with approved academic and commercial research groups [2004.04742]. The database has been shared with **over 30 research groups and companies** since **2014** [2004.04742]. An **open-source Python package** is also mentioned for parsing shared OMI-DB data and providing an API, metadata extraction, and filtering support [2004.04742]. Accordingly, OPTIMAM is a controlled-access research infrastructure rather than an unrestricted public-download dataset.

## 4. OPTIMAM as a lesion-supervised source domain for mammography CAD

In later CAD work, OPTIMAM functioned as a large lesion-supervised training source rather than merely an exam-level label repository. A high-resolution mammography CAD study based on **R-FCN / DCN** used a recently released OPTIMAM dataset consisting of about **78,000 selected digital screening and symptomatic mammograms**, containing approximately **7,500 findings annotated with bounding boxes**, about **6,500** of them cancerous, with **biopsy-determined** ground truth and images mainly from **Hologic** and **GE** machines at around \(4{,}000 \times 5{,}000\) pixels [1902.07323]. The authors state that they selected a detection architecture even though the final task was whole-image classification because they wanted to **“exploit the rich bounding box information in our training dataset (Optimam)”** and improve interpretability [1902.07323].

That study trained on **all OPTIMAM images with findings** and used a **three-class detector** with classes **negative**, **benign**, and **malignant findings**, using the malignant class score for downstream malignancy prediction [1902.07323]. The preprocessing explicitly described for OPTIMAM was manufacturer-specific **intensity normalization via lookup tables**, reflecting the fact that the data span multiple acquisition vendors [1902.07323]. The architectural choices were tightly coupled to OPTIMAM’s image scale and lesion granularity: the model used an input size of **\(2545 \times 2545\)** pixels, batch size **one image per GPU**, and mammography-specific rotation/flip augmentation while avoiding random crops because lesions near image borders could be lost [1902.07323].

The paper did **not** report a dedicated OPTIMAM train/validation/test split or a final held-out OPTIMAM benchmark. Instead, OPTIMAM served as the strongly annotated source domain that enabled transfer to the independent **DREAMS / Group Health** benchmark, where the final detector achieved **AUC \(=0.879\)** for breast-wise detection on **130,000** hidden validation images [1902.07323]. This use of OPTIMAM established a pattern that recurs in subsequent work: lesion-level supervision in OPTIMAM is exploited to train models whose principal quantitative evaluation occurs on separate cohorts.

## 5. OPTIMAM in short-horizon risk prediction and cross-site domain generalization

OPTIMAM has also been used as a multi-site population resource for fixed-horizon risk prediction. In **MamaDino**, the source population came from the **OPTIMAM Mammography Image Database**, specifically women attending routine FFDM screening at **four UK screening services**—**Jarvis, Leicester, Imperial, and St George’s**—between **2010 and 2021**; after filtering, the **training cohort comprised 53,883 women** [2602.13930]. Labels were built from episode outcomes and lesion laterality: episodes with **CI / M / B** outcomes and their screening exams in the preceding **3 years** were labeled positive, normal episodes were labeled negative, and lesion-side OPTIMAM annotations supplied breast-level supervision [2602.13930]. Using **512×512** mammograms, the hybrid **MamaDino + BilateralMixer** model reached **AUC 0.736** on an internal matched case-control test set and **AUC 0.677** on an external out-of-distribution cohort from an unseen OPTIMAM site, Oxford, while the paper states that this matched Mirai despite using **~13× fewer input pixels** [2602.13930]. A plausible implication is that OPTIMAM’s bilateral structure and lesion laterality enable both breast-level and patient-level modeling strategies that would be difficult to construct from image-only repositories.

A different line of work used OPTIMAM as the sole supervised source domain for malignant-versus-benign calcification classification across institutions. That study used **2,994 annotated calcification cases** from OPTIMAM, including **2,187 malignant cases (73.0%)** and **807 benign cases (27.0%)**, and performed **5-fold cross-validation** using only OPTIMAM during training and validation [2607.06549]. The labeled subset was effectively dominated by **Hologic FFDM** screening data, while additional unlabeled OPTIMAM subdomains—**GE FFDM** and **Hologic synthetic** images—were mined for domain adaptation [2607.06549]. The baseline **Swin Transformer V2** achieved **AUC \(0.81 \pm 0.01\)** on OPTIMAM, but external performance fell to **0.68** on both EMBED and Duke; adding the paper’s unsupervised style-transfer adaptation improved AUC from **0.68 to 0.72** on EMBED and from **0.68 to 0.73** on Duke [2607.06549]. This use case shows that OPTIMAM is large and heterogeneous enough to support both supervised lesion classification and adaptation experiments, while also revealing that source-domain strength does not eliminate cross-site domain shift.

## 6. Representativeness, limitations, and scientific position

OPTIMAM’s scale and linkage are accompanied by real-world limitations that matter for experimental design and interpretation. The 2020 systems paper states that lesion locations are **not routinely present** in clinical systems, so only a subset have manual marks; **7,143 lesions** represent only **about 60% of screen-detected cancers** [2004.04742]. Some pathology fields are incomplete in the source systems: **20 screen-detected cancers** and **12 interval cancers** were excluded from the grade table because invasive-status and grade information were missing from the clinical databases [2004.04742]. The cohort assembly includes both exhaustive and sampled components, so years and classes are not uniformly sampled [2004.04742]. The resource reflects the **UK NHS Breast Screening Programme**, where women are typically screened every three years between ages **50 and 70**, with some additional age-trial and high-risk screening groups [2004.04742]. This suggests that OPTIMAM is highly relevant to UK population-screening research but should not be assumed to transfer unchanged to settings with different intervals, age ranges, or case mix.

Within those constraints, OPTIMAM is distinctive because it combines **large scale**, **linked imaging + clinical + pathology data**, **processed and unprocessed images**, **prior mammograms**, **interval cancers**, **partial lesion localization**, **longitudinal updating**, and a governance model that has already supported substantial academic and commercial reuse [2004.04742]. Internal uses reported in the database paper include **virtual clinical trials**, studies of **detector type**, **dose**, and **image processing** on cancer detection, analyses of **cancer characteristics**, and studies of **breast density** in NHSBSP women [2004.04742]. Across later work, OPTIMAM has served as a lesion-annotated detector-training set, a bilateral risk-prediction cohort, and a source domain for multi-site adaptation [1902.07323]. In that sense, OPTIMAM is best understood not as a single benchmark split but as a reusable mammography research infrastructure whose scientific value derives from the joint availability of images, metadata, pathology, follow-up, and partially localized lesions.

Source: https://www.emergentmind.com/topics/optimam