Papers
Topics
Authors
Recent
Search
2000 character limit reached

Credence Calibration Game: Forecasting and Regret

Updated 9 July 2026
  • The Credence Calibration Game is a sequential forecasting framework where agents report probabilities that are evaluated against empirical outcome frequencies.
  • Its methodology partitions rounds by forecast values to compute ℓ1 calibration errors, employing combinatorial tools like Sign-Preservation and SPR to improve error bounds.
  • Recent advances extend the approach to multi-objective, strategic, and LLM settings, integrating calibration with regret minimization and adaptive scoring.

The Credence Calibration Game is a family of sequential forecasting and belief-reporting games in which a predictor, expert, or agent announces a credence—a probability or a distributional forecast—and is evaluated by how closely those announced credences match empirical outcome frequencies on the subsequences where they were issued. In the basic binary online formulation, calibration error after TT rounds is

calerr(T):=∑p∈[0,1]∣mT(p)−p⋅nT(p)∣,\text{calerr}(T) := \sum_{p \in [0,1]} |m_T(p) - p \cdot n_T(p)|,

where nT(p)n_T(p) is the number of times the forecaster predicts pp up to time TT, and mT(p)m_T(p) is the number of those times the outcome was $1$. In this form, a forecaster gives a probability at each step, an adversary chooses binary outcomes, and the forecaster seeks to minimize total calibration error against possibly adversarial choices of outcomes (Dagan et al., 2024).

1. Core formulation and calibration tests

Calibration requires that, for any predicted probability pp, the actual empirical frequency of the event occurring—when the predictor outputs pp—closely match pp. In the binary sequential setting, this yields the calerr(T):=∑p∈[0,1]∣mT(p)−p⋅nT(p)∣,\text{calerr}(T) := \sum_{p \in [0,1]} |m_T(p) - p \cdot n_T(p)|,0-style calibration error above. The same idea extends beyond binary events: in a dynamic expert–decision maker game with state space calerr(T):=∑p∈[0,1]∣mT(p)−p⋅nT(p)∣,\text{calerr}(T) := \sum_{p \in [0,1]} |m_T(p) - p \cdot n_T(p)|,1, a finite calibration test declares forecasts calerr(T):=∑p∈[0,1]∣mT(p)−p⋅nT(p)∣,\text{calerr}(T) := \sum_{p \in [0,1]} |m_T(p) - p \cdot n_T(p)|,2-calibrated when

calerr(T):=∑p∈[0,1]∣mT(p)−p⋅nT(p)∣,\text{calerr}(T) := \sum_{p \in [0,1]} |m_T(p) - p \cdot n_T(p)|,3

where calerr(T):=∑p∈[0,1]∣mT(p)−p⋅nT(p)∣,\text{calerr}(T) := \sum_{p \in [0,1]} |m_T(p) - p \cdot n_T(p)|,4 is the set of periods in which forecast calerr(T):=∑p∈[0,1]∣mT(p)−p⋅nT(p)∣,\text{calerr}(T) := \sum_{p \in [0,1]} |m_T(p) - p \cdot n_T(p)|,5 was issued and calerr(T):=∑p∈[0,1]∣mT(p)−p⋅nT(p)∣,\text{calerr}(T) := \sum_{p \in [0,1]} |m_T(p) - p \cdot n_T(p)|,6 is the empirical distribution of states on those periods. The corresponding asymptotic notion requires the limsup of this quantity to vanish almost surely (Jain et al., 2024).

This formulation makes the game operational. The forecast is not assessed only pointwise, but by conditioning on the subsequences selected by forecast values or forecast bins. In strategic variants, a receiver or tester uses past outcomes to verify those forecasts; if the forecasts pass calibration, the receiver acts as if the forecast is the true state law, while failure can trigger punishment or loss of credibility (Jain et al., 2024).

2. Sequential complexity, lower bounds, and the sign-preservation program

The classical binary online problem has long been organized around two benchmark rates. Foster and Vohra (1998) showed the existence of an algorithm guaranteeing expected calibration error calerr(T):=∑p∈[0,1]∣mT(p)−p⋅nT(p)∣,\text{calerr}(T) := \sum_{p \in [0,1]} |m_T(p) - p \cdot n_T(p)|,7, while anticoncentration arguments yield a lower bound of calerr(T):=∑p∈[0,1]∣mT(p)−p⋅nT(p)∣,\text{calerr}(T) := \sum_{p \in [0,1]} |m_T(p) - p \cdot n_T(p)|,8. Qiao and Valiant (2021) improved the lower bound to calerr(T):=∑p∈[0,1]∣mT(p)−p⋅nT(p)∣,\text{calerr}(T) := \sum_{p \in [0,1]} |m_T(p) - p \cdot n_T(p)|,9 by introducing the Sign-Preservation game, a reduced combinatorial abstraction of the prediction problem. Their reduction theorem translates lower bounds on the game value nT(p)n_T(p)0 into lower bounds on calibration error, and the proof uses the techniques of early stopping and sidestepping to prevent a forecaster from “covering up” earlier miscalibration (Qiao et al., 2020).

A 2024 advance sharpened both the abstraction and the bounds. “Sign Preservation with Reuse” (SPR) generalizes Sign-Preservation by allowing reuse of previously used empty cells under erasure rules; the paper proves a bidirectional relationship between SPR and sequential calibration, rather than only a one-way lower-bound reduction. Concretely, if nT(p)n_T(p)1 for some nT(p)n_T(p)2, then there exists a forecaster with calibration error at most nT(p)n_T(p)3, while if nT(p)n_T(p)4, then for any forecaster, some adversary can ensure calibration error at least nT(p)n_T(p)5. Using this equivalence, the paper gives the first improvement to the nT(p)n_T(p)6 upper bound, obtaining a forecasting algorithm with calibration error nT(p)n_T(p)7 for some absolute nT(p)n_T(p)8, and also proves a lower bound of nT(p)n_T(p)9. That lower bound is obtained by an oblivious adversary, marking the first pp0 calibration lower bound for oblivious adversaries (Dagan et al., 2024).

The algorithmic picture that emerges is explicitly combinatorial. The forecaster simulates multiple parallel SPR games at different scales, tracks “bias” accumulations in discretized probability bins, and “covers up” past deviations by routing current predictions to bins with previously accumulated error. This suggests that the sequential Credence Calibration Game is governed not only by probabilistic concentration, but by a finer preservation-versus-erasure structure encoded by SPR (Dagan et al., 2024).

3. Information structures, randomization, and impossibility phenomena

One major line of work studies calibration in explicit betting or game-theoretic protocols. In Binary Forecasting Game II, Skeptic announces a function pp1, Forecaster announces a probability distribution pp2 over pp3, Reality announces outcome pp4, Forecaster announces pp5 with pp6, a Random Number Generator draws pp7, and Skeptic’s capital is updated by

pp8

Within this framework, Vovk and Shafer showed that randomization can make sequential probability forecasts pass any countable set of well-behaved statistical tests, including calibration tests, but V’yugin showed that this positive result requires forecasts of unrestrictedly increasing degree of accuracy. If forecasts are constrained to a fixed grid with spacing at least pp9, then Reality and Skeptic can force calibration failure, and for either TT0 or TT1,

TT2

The lower bound is therefore controlled by forecast granularity (0808.3746).

A distinct information-theoretic variant arises when uncertainty is represented by a set TT3 of probability distributions and an agent plays against a bookie under the minimax criterion. Two games are central. In the TT4-game, the bookie chooses TT5 before observing TT6; in the TT7-TT8-game, the bookie chooses TT9 after observing mT(p)m_T(p)0. The first game yields an a priori minimax-optimal decision rule, while the second yields an a posteriori minimax-optimal rule based on mT(p)m_T(p)1. Time inconsistency can then be understood as arising because different games are being played against bookies with different information. When mT(p)m_T(p)2 is convex and rectangular, standard conditioning is sharply calibrated; when rectangularity fails, conditioning can fail minimax optimality and sharp calibration can fail (Grunwald et al., 2014).

These results make the information structure part of the definition of the Credence Calibration Game. Calibration is not merely a property of a forecast sequence; it is tied to which player moves first, what the adversary knows, whether forecasts are discrete or arbitrarily precise, and whether updating is evaluated by calibration, minimax loss, or both.

4. Multi-objective and strategic generalizations

In statistical learning, the game is generalized from calibration on one population to simultaneous calibration across many overlapping subpopulations. Discretized multicalibration requires that for every group mT(p)m_T(p)3, credence bin mT(p)m_T(p)4, and class mT(p)m_T(p)5,

mT(p)m_T(p)6

This is formulated as a two-player zero-sum game between a learner, who chooses predictors mT(p)m_T(p)7, and an auditor, who chooses calibration objectives indexed by group, bin, class, and sign. The payoff functions are

mT(p)m_T(p)8

No-regret versus no-regret and no-regret versus best-response dynamics yield near-equilibria and deterministic predictors among iterates. The framework also gives stronger multicalibration conditions that scale with the square-root of group size, removes dependence on mT(p)m_T(p)9 in several settings, and improves $1$0-class multicalibration complexity from $1$1 to $1$2 (Haghtalab et al., 2023).

Strategic interaction provides another extension. Calibrated Stackelberg Games replace the standard assumption that a follower directly observes the leader’s action with the assumption that the follower best-responds to calibrated forecasts about it. The paper introduces adaptive calibration, an any-time notion requiring calibration guarantees on every interval and every bin against adversarial sequences. In the resulting model, the principal can achieve utility converging to the optimum Stackelberg value in both finite and continuous settings, and no higher utility is achievable (Haghtalab et al., 2023).

A related dynamic expert–receiver model studies how an expert sends probabilistic forecasts to maximize utility subject to passing a calibration test. For stationary ergodic processes, the dynamic game reduces to a static persuasion problem: a distribution of forecasts is implementable by a calibrated strategy if and only if it is a mean-preserving contraction of the distribution of conditionals. Against a regret-minimizing decision maker, the expert can achieve the same payoff as under the calibration test, and in some instances can achieve strictly more (Jain et al., 2024).

Taken together, these extensions show that the Credence Calibration Game is not confined to adversarial binary prediction. It functions as a design principle for auditing across groups, for reasoning about best-response dynamics under partial observability, and for constraining strategic communication by statistical credibility.

5. Evaluation criteria, hedging, and connections to regret

A recurrent caution in the literature is that calibration alone is not a complete measure of forecasting quality. “Calibeating” formalizes this point using three quantities: the calibration score

$1$3

the refinement score

$1$4

and the Brier score

$1$5

These satisfy the exact decomposition $1$6. The paper’s central claim is that forecasters should not be tested by calibration score, which can always be made arbitrarily small, but by Brier score, whose refinement component measures how good the sorting into bins with the same forecast is. It then shows that one can calibeat any forecast by an online deterministic procedure, by a stochastic procedure that is itself calibrated, and even simultaneously for multiple procedures (Foster et al., 2022).

Another unifying viewpoint is forecast hedging. Here the forecast is chosen so as to guarantee that the expected track record can only improve. In the binary case, if $1$7 is the gap at forecast $1$8, the one-step change in a calibration score is $1$9, and the hedging condition is to choose pp0 so that pp1 for any pp2. This yields classic calibration results through stochastic, minimax-based procedures and continuous calibration through deterministic, fixed-point-based procedures. The distinction has game-theoretic consequences: stochastic classic calibration yields correlated equilibrium in the long run, while deterministic continuous calibration yields Nash equilibria in the long run (Foster et al., 2022).

A broader evaluative synthesis places a forecaster, a gambler, and nature in a single game. The gambler may choose only gambles pp3 satisfying the availability criterion

pp4

Under intuitive restrictions of this form, calibration and regret emerge as equivalent ways of evaluating forecasts; the paper calls calibration, regret, predictiveness, and randomness the four facets of forecast felicity (Derr et al., 2024). Older work also establishes that calibrated strategies can be obtained by performing strategies that have no internal regret in an auxiliary game, and conversely that a strategy approaching a convex pp5-set can be derived from the construction of a calibrated strategy, including under partial monitoring (Perchet, 2010).

The main misconception addressed by this line of work is that passing a calibration test exhausts the problem of forecast quality. The cited results instead separate calibration from expertise, distinguish discrete from arbitrarily precise randomization, and connect calibration to regret, approachability, and equilibrium selection.

6. LLMs and structured play

Recent work transfers the Credence Calibration Game from classical forecasting to confidence expression by LLMs. One reinforcement-learning approach treats confidence reporting as a betting game: for each question, the model outputs an answer pp6 and an expressed confidence pp7, and receives reward

pp8

The expected reward

pp9

is uniquely maximized at pp0, so the optimal policy under this reward design would result in perfectly calibrated confidence expressions. The implementation generates an answer and then a confidence token, and evaluation uses Expected Calibration Error (ECE), AUROC, and calibration curves (Stangel et al., 4 Mar 2025).

A prompt-based variant explicitly adopts structured play inspired by the Credence Calibration Game. The interaction loop has three stages: pre-game evaluation, a calibration game with repeated rounds of answer-plus-confidence reporting, and post-game evaluation. The model receives natural-language feedback after each round, including total score, accuracy, average confidence, and a qualitative diagnosis such as “You are currently overconfident” or “You are currently underconfident.” Two scoring schemes are used: symmetric scoring, where correct and incorrect answers receive pp1, and exponential scoring, where incorrect high-confidence answers are penalized more steeply. Evaluation on MMLU-Pro and TriviaQA with Llama3.1 and Qwen2.5 models reports consistent improvements in calibration metrics; for example, on TriviaQA, Qwen2.5-7B reduces ECE by pp2 under the exponential game variant (Fang et al., 20 Aug 2025).

These LLM adaptations preserve the core structure of the game: a model reports both an answer and a credence, the credence is scored against correctness, and repeated feedback is used to align stated confidence with observed accuracy. This suggests that the Credence Calibration Game now serves both as a theoretical object in adversarial online learning and as a practical protocol for confidence calibration in generative systems.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Credence Calibration Game.