Papers
Topics
Authors
Recent
Search
2000 character limit reached

ORTHOBO: Orthogonal Bayesian Hyperparameter Optimization

Published 7 May 2026 in cs.LG and cs.AI | (2605.06454v1)

Abstract: Bayesian optimization is widely used for hyperparameter optimization when model evaluations are expensive; however, noisy acquisition estimates can lead to unstable decisions. We identify acquisition estimation noise as a failure mode that was previously overlooked: even when the surrogate model and acquisition target are correctly specified, finite-sample Monte Carlo error can perturb acquisition values. This can, in turn, flip candidate rankings and lead to suboptimal BO decisions. As a remedy, we aim at variance reduction and propose an orthogonal acquisition estimator that subtracts an optimally weighted score-function control variate, which yields an acquisition residual orthogonal to posterior score directions and which thus reduces Monte Carlo variance. We further introduce OrthoBO: a Bayesian optimization framework that combines our orthogonal acquisition estimator with ensemble surrogates and an outer log transformation. We show theoretically that our estimator preserves the target, leads to variance reduction, and improves pairwise ranking stability. We further verify the theoretical properties of OrthoBO through numerical experiments where our framework reduces acquisition estimation variance, stabilizes candidate rankings, and achieves strong performance. We also demonstrate the downstream utility of OrthoBO in hyperparameter optimization for neural network training and fine-tuning.

Summary

  • The paper introduces OrthoBO, an orthogonal score-function control variate for expected improvement that preserves the target while reducing Monte Carlo estimator variance by up to 93.6% in benchmark tests.
  • The method improves candidate-ranking stability and optimization efficiency, reducing pairwise acquisition flips and achieving lower regret than qLogEI and other baselines, especially at small sampling budgets.
  • The full algorithm combines orthogonalized acquisition estimates, surrogate ensembles, and log-transformed expected improvement, delivering gains under model misspecification and in neural-network hyperparameter-tuning tasks.

Motivation and problem statement

Bayesian optimization (BO) is the standard tool for hyperparameter optimization (HPO) when objective evaluations are expensive, but its decisions rest on a quantity that is itself estimated: the acquisition value. The paper "OrthoBO: Orthogonal Bayesian Hyperparameter Optimization" (2605.06454) identifies acquisition estimation noise as a failure mode of BO that the authors argue has been previously overlooked. Even when the surrogate model and the acquisition target are correctly specified, finite-sample Monte Carlo (MC) error in marginalizing over surrogate parameters can perturb acquisition values, flip candidate rankings, and cause BO to evaluate suboptimal configurations. Because BO acts on rankings rather than absolute acquisition values, even small estimation errors can have outsized effects on search trajectories.

The work is positioned as complementary to existing robustness research in BO. Prior approaches address structural surrogate misspecification—kernel misspecification [Bogunovic.2021], unknown hyperparameters [Berkenkamp.2019], unreliable posterior uncertainty [Neiswanger.2021], trust-region methods [Eriksson.2019], ensemble surrogates [Lu.2023, Polyzos.2023]—or improve the acquisition formula and its numerical optimization [Ament.2023]. None directly targets the variance of the estimated acquisition values used for candidate selection. OrthoBO fills this gap by treating acquisition estimation as an MC variance-reduction problem, importing machinery from orthogonal machine learning [Chernozhukov.2018, Foster.2023, Kennedy.2024].

Method: the orthogonal acquisition estimator

The core construction is a control-variate correction to the MC estimator of marginal expected improvement (EI). Given an approximate posterior qm,t(θ)q_{m,t}(\theta) over surrogate parameters θm\theta_m and score function g(θ)=∇θlog⁡qm,t(θ)g(\theta) = \nabla_\theta \log q_{m,t}(\theta), the orthogonalized EI is defined as

EImorth(λ;θ):=EIm(λ;θ)−γm(λ)⊤g(θ),γm(λ):=Σg−1Cov⁡qm,t(g,EIm).\mathrm{EI}_m^{\mathrm{orth}}(\lambda;\theta) := \mathrm{EI}_m(\lambda;\theta) - \gamma_m(\lambda)^\top g(\theta), \qquad \gamma_m(\lambda) := \Sigma_g^{-1}\operatorname{Cov}_{q_{m,t}}(g,\mathrm{EI}_m).

The subtraction is optimally weighted so that the residual is orthogonal to posterior score directions. In practice γm\gamma_m is replaced by an empirical plug-in from the same MC samples; this introduces an O(1/S)O(1/S) in-sample bias dominated by the O(1/S)O(1/\sqrt{S}) sampling noise at the budgets used, and sample-splitting (cross-fitting) would restore exact finite-sample unbiasedness.

A non-trivial implementation detail is that common MC pipelines sample only from the predictive distribution over function values, whereas the score function requires explicit parameter samples; the method therefore draws parameter samples directly and evaluates both EI and gg at the same draws.

The framework provides two instantiations:

  • Gaussian process surrogates: the score function is taken over GP hyperparameters (lengthscales, amplitude, noise variance) under an approximate posterior such as Laplace or variational Gaussian approximations.
  • Tree-based / TPE-style surrogates: since differentiability of qt(θ)q_t(\theta) may fail, the authors generalize to any zero-mean control variate ct(θ)c_t(\theta)—e.g., centered bootstrap statistics of KDE bandwidths and quantile split thresholds—with the same covariance-weighted subtraction.

Theoretical guarantees

Three results anchor the method. Variance reduction (Theorem 1): under square-integrability and nonsingularity of θm\theta_m0, the orthogonalized estimator preserves the marginal EI target exactly and satisfies

θm\theta_m1

which carries over to the MC estimator for any budget θm\theta_m2. Local robustness (Corollary 1): first-order insensitivity to exponential score-tilt perturbations of θm\theta_m3, in the sense of orthogonal ML. Pairwise ranking stability (Proposition 1): via Cantelli's inequality, the probability that the estimated sign of the acquisition gap between two candidates flips is bounded by a decreasing function of the estimator variance, so variance reduction directly reduces ranking errors.

Two caveats are stated plainly. First, these are within-model results: orthogonalization stabilizes estimation conditional on θm\theta_m4 but does not correct structural misspecification of the surrogate family—that role is assigned to ensembling. Second, the plug-in estimate of θm\theta_m5 introduces small-sample bias.

The full OrthoBO algorithm

OrthoBO combines three components with distinct roles: (i) the orthogonal acquisition estimator for variance reduction; (ii) an ensemble of θm\theta_m6 surrogates (different kernels, likelihoods, tree-based models) aggregated via tempered exponentially weighted updates with a minimum weight floor θm\theta_m7, addressing structural misspecification; and (iii) an outer log transformation θm\theta_m8 following Ament et al., improving numerical conditioning during acquisition maximization. Extensions to constrained EI and parallel/batch qEI are given in the supplement, where the generic proposition shows the construction is acquisition-agnostic, requiring only square integrability.

Computational overhead is modest: per-candidate cost adds θm\theta_m9 plus one-time per-iteration preprocessing of g(θ)=∇θlog⁡qm,t(θ)g(\theta) = \nabla_\theta \log q_{m,t}(\theta)0; measured per-step runtimes increase by roughly 0.3–2 seconds relative to qLogEI, which the authors note is negligible relative to expensive objective evaluations.

Empirical findings

Experiments use four synthetic benchmarks (Hartmann6, Ackley8, Michalewicz10, Levy16) against qLogEI, UCB, TuRBO, and Sobol sampling, with 16 replications, g(θ)=∇θlog⁡qm,t(θ)g(\theta) = \nabla_\theta \log q_{m,t}(\theta)1, and g(θ)=∇θlog⁡qm,t(θ)g(\theta) = \nabla_\theta \log q_{m,t}(\theta)2 in the main experiments to isolate the orthogonalization effect.

Variance reduction is substantial. On Michalewicz10 with a Matérn-5/2 ARD kernel, probe variance drops from 9.51×10⁻⁸ to 4.54×10⁻⁸ at g(θ)=∇θlog⁡qm,t(θ)g(\theta) = \nabla_\theta \log q_{m,t}(\theta)3 (−52%) and from 1.67×10⁻⁸ to 0.107×10⁻⁹ at g(θ)=∇θlog⁡qm,t(θ)g(\theta) = \nabla_\theta \log q_{m,t}(\theta)4 (−93.6%); on Levy16, reductions reach −89%. Similar reductions hold for RBF kernels and TPE surrogates.

Ranking stability improves markedly. On Michalewicz10, probe variance falls from 168.6 (qLogEI) to 0.019, top-1 agreement rises from 0.925 to 0.988, and the adjacent-pair flip rate drops from 0.148 to 0.014. On Levy16, regret improves from 82.19 to 58.94.

MC efficiency: across budgets g(θ)=∇θlog⁡qm,t(θ)g(\theta) = \nabla_\theta \log q_{m,t}(\theta)5 on Ackley8, OrthoBO achieves the lowest regret, with gains concentrated in early iterations where MC noise dominates.

Robustness to misspecification is a notable result. Under a strongly misspecified linear kernel, OrthoBO outperforms all baselines on all four functions; on Hartmann6 it reaches qLogEI's final regret level in roughly half the iterations, and on Levy16 it is the only method showing clear improvement. Under weakly fitted hyperparameters (g(θ)=∇θlog⁡qm,t(θ)g(\theta) = \nabla_\theta \log q_{m,t}(\theta)6 on Levy16 with ARD), OrthoBO still attains among the lowest final regrets.

Downstream utility is demonstrated on two real tasks. For CNN training on MNIST/CIFAR10 with 20% injected outlier corruptions, OrthoBO achieves the best final performance. For fine-tuning a ViT-B16 on the industrial WM811K wafer-map dataset (five hyperparameters, F1 objective), OrthoBO improves the best observed validation F1 by up to 20 percentage points over qLogEI—a strong claim supported by only four trials of 20 iterations, which limits statistical confidence. Ensemble experiments (g(θ)=∇θlog⁡qm,t(θ)g(\theta) = \nabla_\theta \log q_{m,t}(\theta)7 kernels) confirm lowest regret and substantially lower probe variance, with weights correctly concentrating on the best-matching Matérn-5/2 kernel.

Limitations and open questions

The paper concedes several limitations. OrthoBO introduces computational overhead, though small relative to expensive objective evaluations. The main-text theory covers only EI-based acquisitions; extension to other acquisition functions "may require acquisition-specific orthogonalization schemes," and while a generic acquisition-level extension is provided for constrained and parallel EI, broader coverage remains unverified empirically. The plug-in estimation of g(θ)=∇θlog⁡qm,t(θ)g(\theta) = \nabla_\theta \log q_{m,t}(\theta)8 sacrifices exact finite-sample unbiasedness unless cross-fitting is adopted. The real-world ViT case study rests on few trials, and the claimed 20-percentage-point improvement should be read with that in mind. Open questions include whether the variance-reduction benefits persist for acquisition families without closed-form marginals beyond nested-MC qEI, and how the temperature and weight-floor hyperparameters of the ensemble aggregation interact with orthogonalization in high-dimensional settings.

Conclusion

OrthoBO reframes acquisition value estimation in BO as an MC variance-reduction problem and solves it with an optimally weighted score-function control variate borrowed from orthogonal ML. The estimator provably preserves the marginal EI target, reduces variance, and improves pairwise ranking stability; empirically it delivers large variance reductions (up to ~93%), more stable rankings, competitive performance under limited MC budgets and misspecified surrogates, and improved downstream HPO for neural network training and ViT fine-tuning. Its principal contribution is identifying and addressing a failure mode—acquisition estimation noise—that prior robustness work in BO had not explicitly targeted.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.