Papers
Topics
Authors
Recent
Search
2000 character limit reached

B-PAC Reasoning: Safety & Efficiency

Updated 6 February 2026
  • B-PAC reasoning is an online decision-making framework that combines inverse propensity scoring, betting martingales, and anytime-valid PAC guarantees to ensure safety and efficiency.
  • It dynamically adjusts the use of high-cost models based on a statistical betting process that monitors uncertainty via thresholding to control risk.
  • Empirical results demonstrate that B-PAC meets stringent performance guarantees and substantially reduces expensive model invocations in non-stationary environments.

Betting Probably Approximately Correct (B-PAC) reasoning refers to a class of online decision-making frameworks that certify performance to within a user-specified tolerance (ε) with high probability (1−α), while optimally minimizing the invocation of an expensive, high-accuracy “thinking model.” The B-PAC paradigm integrates sequential importance sampling, statistical betting martingales, and anytime-valid PAC (Probably Approximately Correct) guarantees. It addresses scenarios where only partial feedback on errors is available and the system must remain robust to non-stationary query streams, yielding practical and theoretically-sound online reasoning mechanisms for large reasoning models and boundedly rational agents (Yu et al., 30 Jan 2026, Oesterheld et al., 2023).

1. Problem Setting: Online Reasoning under Partial Feedback

B-PAC reasoning is motivated by the deployment of two black-box models: a high-accuracy, high-cost “thinking model” f:X→Yf:\mathcal{X}\to\mathcal{Y}, and a low-cost, lower-accuracy “non-thinking model” f′:X→Yf':\mathcal{X}\to\mathcal{Y}. For each time tt, a query XtX_t is drawn from a possibly non-stationary stream. The output f(Xt)f(X_t) is available only when the expensive model is invoked. The decision ft(Xt)∈{f′(Xt),f(Xt)}f_t(X_t)\in\{f'(X_t),f(X_t)\} is determined online by a composite policy.

Crucially, the true response YtY_t is never directly observed; performance loss is measured relative to f(Xt)f(X_t) as ℓ(ft(X),f(X))∈[0,1]\ell(f_t(X),f(X)) \in [0,1], where ℓ\ell is a loss function. The population risk at time f′:X→Yf':\mathcal{X}\to\mathcal{Y}0 under a thresholding policy is defined as

f′:X→Yf':\mathcal{X}\to\mathcal{Y}1

The key objective is to guarantee, with high probability f′:X→Yf':\mathcal{X}\to\mathcal{Y}2, that f′:X→Yf':\mathcal{X}\to\mathcal{Y}3 for all f′:X→Yf':\mathcal{X}\to\mathcal{Y}4 while minimizing invocations of the expensive model.

2. B-PAC Objective: Simultaneous Safety and Efficiency

The dual mandate of B-PAC reasoning is:

  • Safety: Enforce the uniform bound f′:X→Yf':\mathcal{X}\to\mathcal{Y}5 at all times f′:X→Yf':\mathcal{X}\to\mathcal{Y}6 with probability f′:X→Yf':\mathcal{X}\to\mathcal{Y}7 (the “anytime (ε,α)-PAC efficiency guarantee”).
  • Efficiency: Minimize the cumulative fraction of queries routed to f′:X→Yf':\mathcal{X}\to\mathcal{Y}8, measured as

f′:X→Yf':\mathcal{X}\to\mathcal{Y}9

or the equivalent token-level savings.

To achieve both objectives, B-PAC dynamically adapts a threshold tt0 on a scalar uncertainty score tt1 for tt2. If tt3, tt4 is used; otherwise, tt5 is invoked. Raising tt6 expands the usage of tt7, but only when statistical evidence, quantified via a betting process, suffices to certify the performance loss remains within tt8 (Yu et al., 30 Jan 2026).

3. Core Methodology: IPS Risk Estimation and Betting Martingales

The system must estimate the risk of delegating to tt9 under selective feedback. Since XtX_t0 is observed only when XtX_t1 is called, B-PAC employs an inverse propensity scoring (IPS) estimator to correct selection bias. For any threshold XtX_t2, the estimator is:

XtX_t3

where XtX_t4 is the sampling probability (which depends on XtX_t5 and previous thresholds), XtX_t6 indicates whether XtX_t7 was called, and XtX_t8 is the minimal exploration rate.

A “test supermartingale” is constructed as a wealth process XtX_t9. At each round, a nonnegative, predictable bet f(Xt)f(X_t)0 is placed, and

f(Xt)f(X_t)1

with f(Xt)f(X_t)2. Under the null hypothesis that the true risk at threshold f(Xt)f(X_t)3 exceeds f(Xt)f(X_t)4, f(Xt)f(X_t)5 forms a supermartingale, ensuring anytime-valid control via Ville’s inequality (Yu et al., 30 Jan 2026).

4. Thresholding Algorithm and Theoretical Guarantees

A finite grid of candidate thresholds f(Xt)f(X_t)6 is maintained. At each round f(Xt)f(X_t)7:

  1. Compute f(Xt)f(X_t)8 and f(Xt)f(X_t)9.
  2. With probability ft(Xt)∈{f′(Xt),f(Xt)}f_t(X_t)\in\{f'(X_t),f(X_t)\}0 (determined by comparison to ft(Xt)∈{f′(Xt),f(Xt)}f_t(X_t)\in\{f'(X_t),f(X_t)\}1 and the exploration parameter ft(Xt)∈{f′(Xt),f(Xt)}f_t(X_t)\in\{f'(X_t),f(X_t)\}2), decide whether to call ft(Xt)∈{f′(Xt),f(Xt)}f_t(X_t)\in\{f'(X_t),f(X_t)\}3 (i.e., ft(Xt)∈{f′(Xt),f(Xt)}f_t(X_t)\in\{f'(X_t),f(X_t)\}4).
  3. Update IPS estimator ft(Xt)∈{f′(Xt),f(Xt)}f_t(X_t)\in\{f'(X_t),f(X_t)\}5, betting fraction ft(Xt)∈{f′(Xt),f(Xt)}f_t(X_t)\in\{f'(X_t),f(X_t)\}6, and wealth ft(Xt)∈{f′(Xt),f(Xt)}f_t(X_t)\in\{f'(X_t),f(X_t)\}7 for each ft(Xt)∈{f′(Xt),f(Xt)}f_t(X_t)\in\{f'(X_t),f(X_t)\}8.
  4. Set ft(Xt)∈{f′(Xt),f(Xt)}f_t(X_t)\in\{f'(X_t),f(X_t)\}9 to the largest YtY_t0 for which YtY_t1 for all YtY_t2; otherwise, set YtY_t3.

The principal safety guarantee is: with probability YtY_t4, for all YtY_t5, YtY_t6 under i.i.d. queries or certain non-stationary regimes. The adaptive betting strategy, based on online convex optimization (using a projected fraction YtY_t7 incorporating a second-order Taylor surrogate), ensures logarithmic regret YtY_t8 to the best fixed threshold, enabling rapid convergence to optimal efficiency (Yu et al., 30 Jan 2026).

5. Connections to Bounded Inductive Rationality and Hypothesis Betting

B-PAC style reasoning generalizes to settings explored in the theory of bounded inductive rationality (Oesterheld et al., 2023). In this broader context, boundedly rational inductive agents (BRIAs) sequentially choose among a finite set of options YtY_t9 each round, using an internal “betting” process over a countable family of hypotheses f(Xt)f(X_t)0 with computable recommendations and reward promises. BRIAs maintain a policy that never overestimates rewards and test each hypothesis infinitely often. The agent’s “wealth” for each hypothesis is updated via allowance schedules and observed payoffs. Coverage of hypotheses via statistical tests ensures that if any hypothesis promises consistently higher rewards, it is followed sufficiently often.

The general PAC-like guarantee in this setting asserts that, when a valid, high-reward hypothesis exists, the agent’s empirical mean reward converges with high probability to within f(Xt)f(X_t)1 of the best-expected reward, provided payoffs along the hypothesis’ trajectory are sufficiently random (e.g., van-Mises–Wald–Church or Schnorr randomness). In repeated games, pairs of BRIAs can implement any strictly individually rational correlated equilibrium, with empirical plays converging accordingly (Oesterheld et al., 2023).

6. Implementation and Empirical Results

B-PAC reasoning incurs negligible computational overhead compared to large model inference: f(Xt)f(X_t)2 per-round updates for a threshold grid of f(Xt)f(X_t)3 is typical. Exploration probabilities f(Xt)f(X_t)4 are managed by a two-stage schedule: a warm-up phase with f(Xt)f(X_t)5 for the initial f(Xt)f(X_t)6 rounds, then f(Xt)f(X_t)7 for steady-state operation. Asynchronous or sharded implementations of the betting martingale allow for high-throughput deployment (Yu et al., 30 Jan 2026).

Empirical benchmarks on MATH, MMLU-Pro, BIG-Bench Hard (BBH), and Magpie demonstrate substantial computational savings with stringent safety. For example, on Magpie with parameters f(Xt)f(X_t)8, f(Xt)f(X_t)9, B-PAC maintains empirical loss ℓ(ft(X),f(X))∈[0,1]\ell(f_t(X),f(X)) \in [0,1]0 throughout, invokes the expensive model on only ℓ(ft(X),f(X))∈[0,1]\ell(f_t(X),f(X)) \in [0,1]1 of queries, and reserves ℓ(ft(X),f(X))∈[0,1]\ell(f_t(X),f(X)) \in [0,1]2 of tokens for it, outperforming offline PAC thresholds and ablation baselines that either violate safety or are overly conservative. Under non-stationary drifts, B-PAC adaptively tightens the threshold to preserve the safety guarantee, in contrast to fixed offline PAC (Yu et al., 30 Jan 2026).

7. Comparative Analysis, Limitations, and Significance

B-PAC’s integration of IPS estimation, betting supermartingales, and dynamic thresholding provides anytime-valid, model-agnostic, online control of reasoning risks—substantially generalizing classical PAC approaches and outperforming heuristic routing or simple union-bound-based methods. Naive estimators tend to violate safety or incur prohibitive expert-call rates (ℓ(ft(X),f(X))∈[0,1]\ell(f_t(X),f(X)) \in [0,1]3). Classical Hoeffding-style methods are conservative, achieving safety at the expense of efficiency. Heuristic policies such as Chain-of-Draft or NoThinking do not reliably control loss.

A plausible implication is that B-PAC offers a broadly applicable, robust strategy for AI reasoning architectures under uncertainty and limited feedback, and connects deeply to theories of inductive rationality in both single-agent and multi-agent settings. The folk-theorem for BRIAs reinforces that B-PAC-type “betting” approaches are not artifacts of black-box model selection, but arise naturally in settings where optimality and bounded rationality must be simultaneously addressed (Oesterheld et al., 2023).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Betting Probably Approximately Correct (B-PAC) Reasoning.