Component-Wise Attribution of Autoregressive Pretraining Gains

Establish the separate contributions of the autoregressive next-token prediction objective, structured multi-section report templates, abnormality-focused auxiliary outputs, and region-level annotations to downstream long-tailed chest X-ray classification performance through component-wise ablations.

Background

The comparison between MedCLIP-style contrastive pretraining and Med-AR controls the downstream vision architecture and ML-Decoder classification head, but the pretraining recipes differ in supervision, initialization, and compute. Consequently, the observed performance gains cannot be attributed specifically to next-token prediction.

Med-AR uses several additional supervision sources, including structured report sections, abnormality-focused targets, prominence and reasoning fields, and region-level bounding-box descriptions. A component-wise ablation would be needed to determine which elements are responsible for the reported improvements and whether the objective itself provides an independent advantage.

References

Several directions remain open. First, disentangling the contribution of the autoregressive objective itself from its richer supervision---structured multi-section report templates, abnormality-focused auxiliary outputs, and region-level annotations---will require component-wise ablations that we leave to future work.

— Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation  (2609.29156 - Prabhu et al., 24 Sep 2026) in Section 7, “Conclusion and Future Work”