Value of calibrated abstention without an escalation path

Establish whether calibrated abstention has practical value in a deployment without an escalation path, under conditions in which probability-scale confidence is well defined.

Background

The routing pilot measures abstention using percentile-rank scores rather than probabilities. Although the abstaining policy has lower expected calibration error under this rank-based evaluation, the authors interpret the result only as rank-score alignment and not as improved probability calibration.

The practical value of such calibrated abstention remains unresolved, particularly in deployments lacking an escalation mechanism. The paper explicitly states that its experiments do not settle this question.

References

Whether calibrated abstention has value in a deployment without an escalation path is a question for a setting in which a probability-scale confidence is defined; this paper does not settle it.

— Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute  (2609.19942 - Lee et al., 17 Sep 2026) in Section 10, Discussion, subsection "The rank-score calibration effect"