Test the mechanistic explanation involving logit soft-capping

Establish whether tanh-based soft-capping of final logits explains the poor raw-logit energy performance of Gemma 2 2B relative to its NLL performance.

Background

The results report that raw target-logit energy improves performance for several small open-weight models but performs particularly poorly for Gemma 2 2B. The paper notes that Gemma 2 uses tanh-based soft-capping on its final logits, which compresses the raw logit magnitudes used directly by the raw-energy score while leaving normalized NLL comparatively unaffected.

The proposed explanation is explicitly presented as a conjecture rather than as an experimentally established causal account. Verifying it would require isolating the effect of logit soft-capping, for example through controlled model comparisons or calibrated transformations of the logits.

References

Gemma 2 applies tanh-based soft-capping to its final logits , compressing precisely the raw logit magnitude that $S_{Raw}$ reads while leaving the normalized quantity $S_{NLL}$ reads intact. Gemma 4 12B, which does not use the same capping, recovers to 0.263 under the Raw criterion. We offer this explanation as a mechanistic conjecture.

Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking  (2608.17270 - Rajwal et al., 18 Aug 2026) in Section V, Results, subsection “Overall Hypothesis-Identification Performance”