Identify the factors underlying small-model superiority
Identify which training, data, or confidence-calibration factors explain why the one-billion-parameter Llama 3.2 model performs best overall on scientific hypothesis ranking despite its smaller size.
References
Other factors may matter, such as how the model was trained, what data it saw, or how confidently it assigns probabilities to different answers. We did not measure these factors directly, so we cannot say which one explains the result. We see this as an open question for future work.
— Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking
(2608.17270 - Rajwal et al., 18 Aug 2026) in Section VI, Discussion, subsection “Mechanistic Signal May Outperform Explicit Prompting”