Determine whether prompted decoders can close the genre-form performance gap

Determine whether prompted decoder language models can close the performance gap on genre-form classification relative to encoder models.

Background

The paper compares four zero-shot decoder LLMs with supervised encoder models only on 21-class subject classification. The decoders receive no labelled SHELF examples, whereas the encoders use a fitted classifier trained on labelled documents, so the comparison is not a like-for-like evaluation of representation quality.

Genre-form classification is substantially more difficult than subject classification in the reported experiments, and the paper does not evaluate prompted decoders on that task. Consequently, it remains unresolved whether prompting could enable decoder models to match or surpass encoder performance on the harder genre-form label space.

References

We therefore do not know whether prompted decoders would close that gap.

SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking  (2609.03047 - Bommarito, 2 Sep 2026) in Section 6.1, “Reference points and external validity” (Scope and limitations)