Explain the limited utility of high-dimensional query embeddings for confidence estimation
Determine whether the lack of consistent improvements from augmenting uncertainty-score and neighbourhood-statistics features with high-dimensional query embeddings is caused by the dimensionality of the embeddings, which may require substantially more complex models and larger training sets to extract signals predictive of LLM-response correctness.
References
However, we did not observe consistent improvements over models based on uncertainty scores and neighbourhood statistics alone. We conjecture that this is due to the high dimensionality of embedding representations, which may require substantially more complex models and larger training sets to reliably extract signals predictive of correctness.
— Improved Confidence Estimates for Black-Box Large Language Models
(2608.19323 - Mbacke et al., 19 Aug 2026) in Section 3, paragraph “Feature representation”