Final-set F1 benefit of online confidence adaptation

Establish whether online adaptation of the confidence head produces a reliable improvement in final-set F1 beyond its observed candidate-discrimination and tight-budget recall gains.

Background

Online reinforcement-learning adaptation improves the confidence head’s AUROC and produces a statistically supported recall gain at a 10% global return budget on the examined fixed candidate pool. However, the improvement at a 25% budget and the test-oracle final-set F1 difference have uncertainty intervals that include zero.

The paper therefore distinguishes improved local candidate-correctness discrimination from demonstrable improvement in the quality of the selected set. Whether online head adaptation reliably improves final-set F1 remains unresolved.

References

It raises 10\%-budget recall by 1.62 points (95\% interval [0.14, 2.83]); gains at 25\% and in test-oracle F1 remain uncertain.

— Grounding with Confidence: Controllable Generative Video Temporal Grounding  (2609.39883 - Chen et al., 30 Sep 2026) in Section 3.3, “End-to-end grounding and component analysis”