Published 7 May 2026 in cs.LG and cs.AI | (2605.06454v1)
Abstract: Bayesian optimization is widely used for hyperparameter optimization when model evaluations are expensive; however, noisy acquisition estimates can lead to unstable decisions. We identify acquisition estimation noise as a failure mode that was previously overlooked: even when the surrogate model and acquisition target are correctly specified, finite-sample Monte Carlo error can perturb acquisition values. This can, in turn, flip candidate rankings and lead to suboptimal BO decisions. As a remedy, we aim at variance reduction and propose an orthogonal acquisition estimator that subtracts an optimally weighted score-function control variate, which yields an acquisition residual orthogonal to posterior score directions and which thus reduces Monte Carlo variance. We further introduce OrthoBO: a Bayesian optimization framework that combines our orthogonal acquisition estimator with ensemble surrogates and an outer log transformation. We show theoretically that our estimator preserves the target, leads to variance reduction, and improves pairwise ranking stability. We further verify the theoretical properties of OrthoBO through numerical experiments where our framework reduces acquisition estimation variance, stabilizes candidate rankings, and achieves strong performance. We also demonstrate the downstream utility of OrthoBO in hyperparameter optimization for neural network training and fine-tuning.
The paper introduces OrthoBO, an orthogonal score-function control variate for expected improvement that preserves the target while reducing Monte Carlo estimator variance by up to 93.6% in benchmark tests.
The method improves candidate-ranking stability and optimization efficiency, reducing pairwise acquisition flips and achieving lower regret than qLogEI and other baselines, especially at small sampling budgets.
The full algorithm combines orthogonalized acquisition estimates, surrogate ensembles, and log-transformed expected improvement, delivering gains under model misspecification and in neural-network hyperparameter-tuning tasks.
Motivation and problem statement
Bayesian optimization (BO) is the standard tool for hyperparameter optimization (HPO) when objective evaluations are expensive, but its decisions rest on a quantity that is itself estimated: the acquisition value. The paper "OrthoBO: Orthogonal Bayesian Hyperparameter Optimization" (2605.06454) identifies acquisition estimation noise as a failure mode of BO that the authors argue has been previously overlooked. Even when the surrogate model and the acquisition target are correctly specified, finite-sample Monte Carlo (MC) error in marginalizing over surrogate parameters can perturb acquisition values, flip candidate rankings, and cause BO to evaluate suboptimal configurations. Because BO acts on rankings rather than absolute acquisition values, even small estimation errors can have outsized effects on search trajectories.
The work is positioned as complementary to existing robustness research in BO. Prior approaches address structural surrogate misspecification—kernel misspecification [Bogunovic.2021], unknown hyperparameters [Berkenkamp.2019], unreliable posterior uncertainty [Neiswanger.2021], trust-region methods [Eriksson.2019], ensemble surrogates [Lu.2023, Polyzos.2023]—or improve the acquisition formula and its numerical optimization [Ament.2023]. None directly targets the variance of the estimated acquisition values used for candidate selection. OrthoBO fills this gap by treating acquisition estimation as an MC variance-reduction problem, importing machinery from orthogonal machine learning [Chernozhukov.2018, Foster.2023, Kennedy.2024].
Method: the orthogonal acquisition estimator
The core construction is a control-variate correction to the MC estimator of marginal expected improvement (EI). Given an approximate posterior qm,t(θ) over surrogate parameters θm and score function g(θ)=∇θlogqm,t(θ), the orthogonalized EI is defined as
The subtraction is optimally weighted so that the residual is orthogonal to posterior score directions. In practice γm is replaced by an empirical plug-in from the same MC samples; this introduces an O(1/S) in-sample bias dominated by the O(1/S) sampling noise at the budgets used, and sample-splitting (cross-fitting) would restore exact finite-sample unbiasedness.
A non-trivial implementation detail is that common MC pipelines sample only from the predictive distribution over function values, whereas the score function requires explicit parameter samples; the method therefore draws parameter samples directly and evaluates both EI and g at the same draws.
The framework provides two instantiations:
Gaussian process surrogates: the score function is taken over GP hyperparameters (lengthscales, amplitude, noise variance) under an approximate posterior such as Laplace or variational Gaussian approximations.
Tree-based / TPE-style surrogates: since differentiability of qt(θ) may fail, the authors generalize to any zero-mean control variate ct(θ)—e.g., centered bootstrap statistics of KDE bandwidths and quantile split thresholds—with the same covariance-weighted subtraction.
Theoretical guarantees
Three results anchor the method. Variance reduction (Theorem 1): under square-integrability and nonsingularity of θm0, the orthogonalized estimator preserves the marginal EI target exactly and satisfies
θm1
which carries over to the MC estimator for any budget θm2. Local robustness (Corollary 1): first-order insensitivity to exponential score-tilt perturbations of θm3, in the sense of orthogonal ML. Pairwise ranking stability (Proposition 1): via Cantelli's inequality, the probability that the estimated sign of the acquisition gap between two candidates flips is bounded by a decreasing function of the estimator variance, so variance reduction directly reduces ranking errors.
Two caveats are stated plainly. First, these are within-model results: orthogonalization stabilizes estimation conditional on θm4 but does not correct structural misspecification of the surrogate family—that role is assigned to ensembling. Second, the plug-in estimate of θm5 introduces small-sample bias.
The full OrthoBO algorithm
OrthoBO combines three components with distinct roles: (i) the orthogonal acquisition estimator for variance reduction; (ii) an ensemble of θm6 surrogates (different kernels, likelihoods, tree-based models) aggregated via tempered exponentially weighted updates with a minimum weight floor θm7, addressing structural misspecification; and (iii) an outer log transformation θm8 following Ament et al., improving numerical conditioning during acquisition maximization. Extensions to constrained EI and parallel/batch qEI are given in the supplement, where the generic proposition shows the construction is acquisition-agnostic, requiring only square integrability.
Computational overhead is modest: per-candidate cost adds θm9 plus one-time per-iteration preprocessing of g(θ)=∇θlogqm,t(θ)0; measured per-step runtimes increase by roughly 0.3–2 seconds relative to qLogEI, which the authors note is negligible relative to expensive objective evaluations.
Empirical findings
Experiments use four synthetic benchmarks (Hartmann6, Ackley8, Michalewicz10, Levy16) against qLogEI, UCB, TuRBO, and Sobol sampling, with 16 replications, g(θ)=∇θlogqm,t(θ)1, and g(θ)=∇θlogqm,t(θ)2 in the main experiments to isolate the orthogonalization effect.
Variance reduction is substantial. On Michalewicz10 with a Matérn-5/2 ARD kernel, probe variance drops from 9.51×10⁻⁸ to 4.54×10⁻⁸ at g(θ)=∇θlogqm,t(θ)3 (−52%) and from 1.67×10⁻⁸ to 0.107×10⁻⁹ at g(θ)=∇θlogqm,t(θ)4 (−93.6%); on Levy16, reductions reach −89%. Similar reductions hold for RBF kernels and TPE surrogates.
Ranking stability improves markedly. On Michalewicz10, probe variance falls from 168.6 (qLogEI) to 0.019, top-1 agreement rises from 0.925 to 0.988, and the adjacent-pair flip rate drops from 0.148 to 0.014. On Levy16, regret improves from 82.19 to 58.94.
MC efficiency: across budgets g(θ)=∇θlogqm,t(θ)5 on Ackley8, OrthoBO achieves the lowest regret, with gains concentrated in early iterations where MC noise dominates.
Robustness to misspecification is a notable result. Under a strongly misspecified linear kernel, OrthoBO outperforms all baselines on all four functions; on Hartmann6 it reaches qLogEI's final regret level in roughly half the iterations, and on Levy16 it is the only method showing clear improvement. Under weakly fitted hyperparameters (g(θ)=∇θlogqm,t(θ)6 on Levy16 with ARD), OrthoBO still attains among the lowest final regrets.
Downstream utility is demonstrated on two real tasks. For CNN training on MNIST/CIFAR10 with 20% injected outlier corruptions, OrthoBO achieves the best final performance. For fine-tuning a ViT-B16 on the industrial WM811K wafer-map dataset (five hyperparameters, F1 objective), OrthoBO improves the best observed validation F1 by up to 20 percentage points over qLogEI—a strong claim supported by only four trials of 20 iterations, which limits statistical confidence. Ensemble experiments (g(θ)=∇θlogqm,t(θ)7 kernels) confirm lowest regret and substantially lower probe variance, with weights correctly concentrating on the best-matching Matérn-5/2 kernel.
Limitations and open questions
The paper concedes several limitations. OrthoBO introduces computational overhead, though small relative to expensive objective evaluations. The main-text theory covers only EI-based acquisitions; extension to other acquisition functions "may require acquisition-specific orthogonalization schemes," and while a generic acquisition-level extension is provided for constrained and parallel EI, broader coverage remains unverified empirically. The plug-in estimation of g(θ)=∇θlogqm,t(θ)8 sacrifices exact finite-sample unbiasedness unless cross-fitting is adopted. The real-world ViT case study rests on few trials, and the claimed 20-percentage-point improvement should be read with that in mind. Open questions include whether the variance-reduction benefits persist for acquisition families without closed-form marginals beyond nested-MC qEI, and how the temperature and weight-floor hyperparameters of the ensemble aggregation interact with orthogonalization in high-dimensional settings.
Conclusion
OrthoBO reframes acquisition value estimation in BO as an MC variance-reduction problem and solves it with an optimally weighted score-function control variate borrowed from orthogonal ML. The estimator provably preserves the marginal EI target, reduces variance, and improves pairwise ranking stability; empirically it delivers large variance reductions (up to ~93%), more stable rankings, competitive performance under limited MC budgets and misspecified surrogates, and improved downstream HPO for neural network training and ViT fine-tuning. Its principal contribution is identifying and addressing a failure mode—acquisition estimation noise—that prior robustness work in BO had not explicitly targeted.
“Emergent Mind helps me see which AI papers have caught fire online.”
Philip
Creator, AI Explained on YouTube
Sign up for free to explore the frontiers of research
Discover trending papers, chat with arXiv, and track the latest research shaping the future of science and technology.Discover trending papers, chat with arXiv, and more.