Establish whether sample-based configuration rankings transfer to complete corpora

Determine whether a configuration ranked lower on stratified readability-assessment samples would outperform the selected configuration when evaluated on the complete corpora.

Background

The experimental configurations are ranked using stratified samples of at most 567 documents from each corpus, and only the highest-ranked configuration is subsequently evaluated on the complete corpora. Consequently, the study does not verify whether the sample-based ranking remains valid at full scale.

The mean Spearman correlation declines from 0.772 on the sampled datasets to 0.694 on the complete corpora. The authors therefore leave unresolved whether another configuration, ranked lower on the samples, would transfer better to the full datasets.

References

Mean $\rho$ falls from $0.772$ on the samples to $0.694$ on the complete corpora, and we cannot rule out that a configuration ranked lower on the samples would transfer better.

— Assessing Readability with LLMs: The Role of Reasoning and Few-Shot Prompting  (2609.24650 - Thieffry et al., 21 Sep 2026) in Section 6, Limitations, item 3