Paper: OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers
Listen to this article.
Audio is available for 30 days and will be removed automatically.
Problem
Selecting an optimizer for training large-scale machine learning models has become surprisingly complex. With over one hundred available methods, researchers and engineers are facing a fragmented landscape. The choice isn’t just about performance; it’s a system-level design decision that must balance computational resources, memory constraints, the effort required for tuning, and the specific requirements of different tasks.