Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods
Pulling back the Fisher metric yields a local sensitivity tensor that bounds acquisition gradients and explains high-dimensional BO failures, leading to the FITR trust-region method that replaces lengthscale heuristics with local Fisher weights.
Published 2026Sydney Poster Session 3 · Wed, Dec 9, 10:00 AM–1:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

Only vote on papers you've read. Sign in with GitHub to vote.
A powerful Fisher pullback unifies high-dimensional BO vanishing gradients and heuristics like RAASP, though FITR's gains remain task-dependent and its geometric weight formula is underspecified.
Abstract
We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of reparameterizable acquisition functions. This view explains vanishing-gradient behavior in high-dimensional BO and provides a common interpretation of heuristics such as RAASP and dimension-scaled lengthscales. Building on this analysis, we propose FITR, a trust-region-based BO method that replaces lengthscale-based scaling by local pullback-Fisher weights. FITR is not restricted to GP kernels with explicit lengthscales. On GP benchmarks with an SE kernel, experiments show competitive performance using FITR. The proposed method also easily generalizes to non-isotropic surrogates, although the gains are more task-dependent in that setting.