Analyze APX–APS mismatch cases qualitatively

Perform a systematic qualitative analysis of cases in which Accuracy on Prompt Execution and Accuracy on Perplexity Score disagree, in order to characterize the sources and patterns of the mismatch.

Background

The paper reports a robust aggregate divergence between prompted answer selection, measured by Accuracy on Prompt Execution (APX), and declarative-statement likelihood ranking, measured by Accuracy on Perplexity Score (APS). However, the aggregate results do not explain why individual questions produce different outcomes under the two protocols. A systematic qualitative examination of disagreement cases is left unresolved and could reveal whether mismatches arise from answer-selection behavior, surface-form effects, semantic distortions, or other properties of the evaluated questions and statements.

References

We also leave a systematic qualitative analysis of APX--APS mismatch cases to future work.

— Likelihood Ranking doesn't Scale Like Prompting in LLMs  (2609.29390 - Bondielli et al., 24 Sep 2026) in Limitations, final paragraph of Section 6