Transfer of MMLU structural sensitivities to closed frontier models

Determine whether the structural sensitivities identified by the text-based audit of MMLU extend to closed frontier language models.

Background

The study evaluates open-weights LLMs and finds that MMLU difficulty reflects separable retrieval and reasoning-related structures. Its conclusions therefore may not generalize to proprietary or otherwise closed frontier systems, whose training data, architectures, and capabilities may differ from those of the evaluated models. The paper explicitly leaves unresolved whether the observed sensitivities are also present in that broader model population.

References

This evaluation also targets open-weights models, and whether these sensitivities extend to closed frontier models remains open.

What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores  (2609.09372 - Paquin et al., 8 Sep 2026) in Section Limitations, subsection “Variance explained and scope”