Generality of the calibration–correctness trade-off

Determine whether the trade-off between high item-level correctness and poor calibration in individualized answer simulation is specific to the Qwen3-1.7B implementation and its retrieval configuration or reflects a general limitation of individual-text-corpus-based answer simulation.

Background

The main retrieval-augmented-generation analysis evaluates only the Qwen3-1.7B architecture with a single retrieval configuration. Although the model achieves relatively strong factual correctness, it assigns poorly calibrated probabilities to the answers selected by individual participants.

Because alternative model sizes, calibration methods, and retrieval strategies were not compared in the main setting, the paper leaves unresolved whether this performance trade-off is an artifact of the particular implementation or an inherent limitation of simulating individual answer distributions from individual text corpora.

References

Because no alternative model sizes, calibration procedures, or retrieval strategies were compared, it remains unclear whether the observed trade-off is specific to the present implementation or reflects a more general limitation of IC-based answer simulation.

Individual Text Corpora Predict User-Specific Knowledge: Benchmarks of Individualized Knowledge Simulation  (2609.08532 - Wigbels et al., 8 Sep 2026) in Limitations, fifth limitation