Scaling subspace Bayesian inference to foundation-model-scale reward models

Determine whether neural-network subspace methods can effectively apply Bayesian inference to foundation-model-scale reward models, with the aim of retaining tractable uncertainty quantification at substantially larger model sizes.

Background

The paper introduces PreferenceEKF, which performs extended Kalman filtering in a low-dimensional subspace of a neural-network reward model to reduce the computational and memory costs of maintaining parameter uncertainty. The experiments demonstrate scalability for the comparatively small reward models evaluated in the D4RL and V-D4RL settings.

The authors explicitly identify the unresolved question of whether this subspace-based approach will remain effective for foundation-model-scale reward models. This concerns the applicability of their uncertainty-representation and Bayesian-inference strategy to much larger models, rather than merely a proposed future evaluation.

References

While we found subspace methods to be an effective tool for scaling Bayesian filtering methods for neural network training, it is unclear whether this approach will be effective for applying Bayesian methods to foundation model-scale reward models (Mahan et al., 2024; Zhang et al., 2024).

Subspace Inference Enables Efficient Active Reward Learning from Preferences  (2609.04066 - Zhou et al., 3 Sep 2026) in Section 6, “Limitations and future work”