AI/ bayesian-optimization · information-geometry · gaussian-processes · machine-learning

A Geometry Fix for Bayesian Optimization's Vanishing Gradients

Researchers use Fisher information geometry to explain vanishing gradients in high-dimensional Bayesian optimization, proposing a trust-region fix called FITR.

A new paper diagnoses why Bayesian optimization slows to a crawl in high dimensions - and offers a geometric fix.

Researchers examined Bayesian optimization through information geometry, pulling the Fisher information metric back through the surrogate model's posterior map. This produces a local sensitivity tensor over the input space that bounds the gradient of common acquisition functions, explaining the vanishing-gradient problem that plagues high-dimensional BO. The framework also unifies two existing heuristics, RAASP and dimension-scaled lengthscales, showing they are really the same trick viewed from different angles. From that foundation, the authors built FITR, a trust-region BO method that swaps lengthscale-based scaling for local pullback-Fisher weights.

BO is the standard tool for tuning hyperparameters and other expensive-to-evaluate systems, and it has long been known to stall past a few dozen dimensions. Explaining that failure geometrically, rather than patching it with another heuristic, gives practitioners a principled reason to trust - or distrust - their scaling choices. It also frees BO from a GP-kernel straitjacket: FITR does not require an explicit lengthscale, so it can extend to non-isotropic surrogates.

On standard GP benchmarks with a squared-exponential kernel, FITR performs competitively, though the paper is candid that gains outside that setting are task-dependent - not the kind of caveat you get in a launch post dressed up as a breakthrough.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →