Forecasting accuracy of large language models for future events
Determine the forecasting accuracy of large language models, specifically OpenAI’s GPT-3.5 and GPT-4, when used as devices to predict future events that lie beyond their training data.
References
But how accurate they are is unknown in part because these new technologies seem poorly understood even by its creators.
— Can Base ChatGPT be Used for Forecasting without Additional Optimization?
(2404.07396 - Pham et al., 2024) in Introduction
We therefore cannot estimate the share of the gap attributable to breadth or interpret a residual as pure forecasting ability. Such attribution requires a controlled specificity intervention or better-matched samples.
— Can Large Language Models Forecast What Researchers Study Next?
(2609.00747 - Li et al., 1 Sep 2026) in Section 4.3, paragraph “What remains unresolved”; Appendix, Section “Outcome-Blind Generality Analysis”