Usefulness of superforecasters for AI risk estimation

Determine whether well-calibrated generalist forecasters or superforecasters can reliably estimate the probability that a large language model capability produces meaningful operational uplift for a specific step in a cyberattack.

Background

The paper considers human domain experts, human forecasting experts, and LLM-simulated experts as possible sources of probabilistic estimates. Superforecasters could potentially supplement scarce technical experts, but it is uncertain whether their general forecasting ability transfers to highly technical and novel AI-risk questions. The concrete example identified by the authors concerns estimating whether a particular LLM capability improves performance at a specific stage of a cyberattack.

References

The question of whether superforecasters can usefully estimate, for example, the probability that a given LLM capability translates into meaningful uplift for a specific step in a cyberattack remains open.

Open Problems in AI Risk Modeling: Insights from a Workshop on the Technical Foundations of AI Risk Modeling  (2609.03178 - Jackson et al., 2 Sep 2026) in Section 4.3.1, “Elicitation Methods”