Reconcile the FHM test-set denominator discrepancy

Reconcile the discrepancy between the displayed FHM test-seen count of 1,000 examples and the saved paired-analysis artifact containing 931 paired predictions for the Gemma-3-12B rank-32 bilinear readout evaluation.

Background

The FHM cross-modal experiment evaluates a Gemma-3-12B rank-32 bilinear readout against the native model. The reported test-seen macro-F1 comparison uses a displayed denominator of 1,000 examples, while the underlying saved paired-analysis artifact contains predictions for only 931 paired examples. Because the reported comparison and paired statistical analyses may depend on the evaluation denominator, the inconsistency must be resolved to establish the precise provenance and reproducibility of the test results.

References

The displayed test-seen count is $n = 1000$, following the manuscript's current reporting convention; the saved paired-analysis artifact underlying the unchanged scores contains 931 paired predictions. This denominator discrepancy remains to be reconciled.

Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection  (2609.18860 - Koushik et al., 16 Sep 2026) in Figure 6 caption, Appendix H (FHM Low-Rank Cross-Modal Readout Details)