Good Papers

Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods

Pulling back the Fisher metric yields a local sensitivity tensor that bounds acquisition gradients and explains high-dimensional BO failures, leading to the FITR trust-region method that replaces lengthscale heuristics with local Fisher weights.

Saksham Kiroriwal, Julius Pfrommer, Jürgen Beyerer

Published 2026Sydney Poster Session 3 · Wed, Dec 9, 10:00 AM–1:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

72%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel8/20reviewers recommend it
lenient 2/5
medium 5/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
A powerful Fisher pullback unifies high-dimensional BO vanishing gradients and heuristics like RAASP, though FITR's gains remain task-dependent and its geometric weight formula is underspecified.

Abstract

We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of reparameterizable acquisition functions. This view explains vanishing-gradient behavior in high-dimensional BO and provides a common interpretation of heuristics such as RAASP and dimension-scaled lengthscales. Building on this analysis, we propose FITR, a trust-region-based BO method that replaces lengthscale-based scaling by local pullback-Fisher weights. FITR is not restricted to GP kernels with explicit lengthscales. On GP benchmarks with an SE kernel, experiments show competitive performance using FITR. The proposed method also easily generalizes to non-isotropic surrogates, although the gains are more task-dependent in that setting.