Identify the factors underlying small-model superiority

Identify which training, data, or confidence-calibration factors explain why the one-billion-parameter Llama 3.2 model performs best overall on scientific hypothesis ranking despite its smaller size.

Background

The strongest configuration in the experiments is Llama 3.2 1B scored with raw target-logit energy, while larger models do not consistently perform better. The authors suggest that training procedure, training data, and the model’s confidence assignment may influence performance, but these factors were not measured directly.

The paper therefore leaves unresolved which properties account for the observed relationship between model size and hypothesis-ranking performance. Resolving this question would clarify whether the result reflects architecture, data exposure, calibration, or another model characteristic.

References

Other factors may matter, such as how the model was trained, what data it saw, or how confidently it assigns probabilities to different answers. We did not measure these factors directly, so we cannot say which one explains the result. We see this as an open question for future work.

Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking  (2608.17270 - Rajwal et al., 18 Aug 2026) in Section VI, Discussion, subsection “Mechanistic Signal May Outperform Explicit Prompting”