Stochastic Weight Averaging: Finding Wider Optima for Better Generalization
Introduction Deep neural networks are typically trained by optimizing a loss function with SGD and a decaying learning rate until convergence. The standard approach works well, but recent work on loss surface geometry suggests we can do better. Resea...
Dec 16, 20257 min read