Determine the causal sources of cross-cohort performance differences

Determine the individual causal contributions of acquisition variation, disease-spectrum variation, and label variation to the observed differences between chest X-ray tuberculosis evaluation cohorts.

Background

The audit compares Montgomery, Shenzhen, TBX11K, and VinDr-CXR, but the cohorts differ simultaneously in image acquisition, disease spectrum, and labeling practices. The paper explicitly states that the study cannot identify how much each factor contributes to the observed changes in discrimination and operating performance.

Resolving these contributions would clarify whether the reported portability failures arise primarily from imaging-site differences, variation in the clinical composition of non-tuberculosis cases, inconsistencies in tuberculosis labels, or interactions among these factors.

References

Differences between cohorts combine acquisition, spectrum and label variation; their individual causal contributions remain unidentified.

— Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening  (2609.21763 - Akhtar et al., 18 Sep 2026) in Section 3.6, “Limitations and priorities for further work”

These studies could determine whether the identified dependencies remain consequential under actual screening conditions.

— Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening  (2609.21763 - Akhtar et al., 18 Sep 2026) in Section 3.6, “Limitations and priorities for further work”