Good Papers

Bidirectional Information Flow (BIF) - A Sample Efficient Hierarchical Gaussian Process for Bayesian Optimization

Bidirectional Information Flow enables continuous two-way communication in hierarchical Gaussian processes for Bayesian optimization, improving sample efficiency, training robustness, and modular subtask reuse while significantly outperforming unidirectional and vanilla methods.

Juan D. Guerra, Thomas Garbay, Numa Dancause, Guillaume Lajoie, Marco Bonizzato

Published 2026Sydney Poster Session 2 · Tue, Dec 8, 5:00 PM–8:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

88%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel15/20reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Hierarchical Gaussian Process (H-GP) models divide problems into different subtasks, allowing different components to address each part, making them well-suited for problems with inherent compositional structure. However, existing H-GP frameworks typically employ one-way information sharing - either top-down or bottom-up - which limits sample efficiency and slows convergence. We propose Bidirectional Information Flow (BIF), which establishes continuous two-way communication. BIF retains the modular structure of hierarchical models-the parent conditions its own posterior on child summaries, treating them as structured priors-while introducing top-down feedback to softly decompose environment observations from the parent into sub-responses. This mutual exchange improves sample efficiency, enables robust training, and allows modular reuse of learned subtask models. We prove analytically that the regret of a GP with a learned kernel scales linearly with the mismatch to the true kernel, tightening in the hierarchical case to the sum of child-level errors. Ablation shows that removing the downward pathway collapses child $R^2$ by up to 58%. Across synthetic, neurostimulation, and HPO benchmarks, BIF achieves up to $4\times$ higher parent $R^2$ and $\sim 100\%$ AUC improvement over vanilla GPBO, and outscores all hierarchical state-of-the-art methods on child $R^2$ given the correct acquisition function, while supporting modular child transfer to novel composite tasks.