Hubness-aware modality-gap correction

Develop modality-gap correction objectives that reduce prediction concentration while preserving the benefits of image–text alignment, thereby making gap reduction explicitly hubness-aware.

Background

The paper shows that reducing the average modality gap through Linear correction can distort class-wise decision margins, concentrate predictions on a small number of classes, and degrade zero-shot classification accuracy. CSLS and modality-wise centering provide only partial or diagnostic mitigation, and the authors state that neither constitutes a general solution. The unresolved research direction is therefore to design correction objectives that directly control prediction-level hubness without sacrificing the alignment improvements sought by modality-gap reduction.

References

We present this as a candidate correlate, not a causal explanation (the correlation reverses on EuroSAT under the generic template; Appendix~\ref{app:generic_config}); whether uniformity can be constrained during correction or post-training is future work.

When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP  (2609.01103 - Sato et al., 1 Sep 2026) in Section 5, paragraph “A Candidate Geometric Correlate”

Making gap reduction itself hubness-aware is left to future work.

When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP  (2609.01103 - Sato et al., 1 Sep 2026) in Section 4.4, “Interpreting and Mitigating Prediction Concentration” (labelled `subsec:mitigation`)

Thus, the metric used to quantify the gap is not identical to the scoring rule used for prediction. This mismatch suggests that future work should consider gap measures that are more directly tied to the geometry of downstream decision rules.

When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP  (2609.01103 - Sato et al., 1 Sep 2026) in Section 5, paragraph “Is the Average Modality Gap the Right Quantity?”