Cross-dataset consistency of confidence-based ranking

Characterize whether confidence-based selection consistently improves single-answer temporal-grounding ranking across Charades, ActivityNet, and QVHighlights.

Background

The paper compares returning the first generated interval with returning the highest-confidence interval on matched subsets of three single-interval datasets. Confidence improves mean temporal IoU on all three datasets at the point-estimate level, but the bootstrap intervals for ActivityNet and QVHighlights include zero, unlike the interval for Charades.

Consequently, the evidence does not establish that the ranking benefit generalizes across datasets. The unresolved issue is whether the confidence head provides a robust cross-dataset ranking improvement rather than a dataset-specific gain.

References

This supports within-query ranking on Charades while leaving its consistency across datasets unresolved.

— Grounding with Confidence: Controllable Generative Video Temporal Grounding  (2609.39883 - Chen et al., 30 Sep 2026) in Section 3.4, “Ranking and rejection on single-interval datasets”