Jointly identifying label priors and offsets without supervision

Determine how to identify both the true label prior and the per-label additive offsets from the same unlabeled score matrix without additional supervision, particularly when the deployment label distribution may be imbalanced or unknown.

Background

The paper’s diagnostic correction assumes a supplied label prior and evaluates balanced datasets, so the method is tested only in a regime where the prior is known and uniform. In practical deployments, the label distribution may be imbalanced or unavailable, creating a joint identification problem: the observed score matrix must simultaneously reveal the underlying label prior and the global offsets that correct structural scoring biases.

The paper explicitly leaves unresolved whether these two quantities can be recovered from unlabeled scores alone. Resolving this problem would extend the prior-conditioned additive correction beyond balanced benchmark settings and remove its dependence on an externally specified operating prior.

References

Identifying both the prior and the offsets from the same unlabeled score matrix without additional supervision remains an open challenge.

— Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores  (2608.31068 - Yan et al., 31 Aug 2026) in Section* Limitations