Source of EA-Qwen’s calibration advantage

Determine what produces EA-Qwen’s calibration advantage over Direct-Qwen, which persists after matching the search configurations on search difficulty and cannot be explained by the amount of graded search history available to the operator.

Background

The paper reports that EA-Qwen, which places the Qwen-3 14B LLM inside an evolutionary search loop, has better confidence calibration than Direct-Qwen, which uses the same model in sequential search. Matching the two configurations on search difficulty explains only part of the calibration difference.

The remaining calibration advantage is not explained by the evolutionary configuration’s access to more graded history: Direct-Qwen sees the complete ranked guess history, whereas EA-Qwen sees only its parent’s tried words without their ranks. The authors explicitly leave the mechanism producing this persistent difference unresolved. This is therefore a concrete open explanatory problem concerning the causal source of the observed calibration effect.

References

The source of EA-Qwen's calibration advantage remains unresolved.

Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search  (2609.00652 - Pan et al., 1 Sep 2026) in Appendix, Section Additional Discussion and Future Work