SoftServe: a quasi-Newton optimizer that keeps curvature estimates positive and scales to huge neural nets
This paper introduces SoftServe, a family of optimization methods for training deep neural networks that adapts quasi-Newton ideas to the non-convex, large-scale setting of deep learning. The main idea is to produce curvature estimates that are guaranteed to be positive definite (so they point downhill) even when the true second-derivative information contains negative directions. SoftServe does this without costly line searches or ad hoc fixes that are hard to use with stochastic mini-batches.
Quasi-Newton methods try to speed up optimization by using an estimate of the inverse Hessian, a matrix that captures curvature of the loss function. A common way to build that estimate enforces the secant equation, which ties changes in parameters to changes in gradients. But for non-convex objectives the secant equation can force the estimate to have negative eigenvalues, which can make an optimizer move uphill. Line searches can avoid that in classical optimization, but they need full-batch loss evaluations and so are impractical for modern stochastic training.
SoftServe builds on a recent idea that relaxes the hard secant constraint into a soft penalty. That variational formulation keeps the update well defined and positive definite even when curvature is negative. The authors make this idea practical for deep learning in two ways. First, they restrict the inverse-Hessian estimate to structured families that are cheap to store and update: a diagonal family (called -Diag) and a Kronecker-factored family (called -Kron). Second, they replace expensive matrix decompositions with a stable Newton–Schulz iteration, which uses matrix multiplications that run well on GPUs (graphics processing units).
Why this matters: the two structured variants bring memory and compute costs down to the same order as popular adaptive optimizers. -Diag uses O(D) cost, comparable to Adam. -Kron uses costs like Shampoo or KFAC for a parameter block. The paper reports that SoftServe performs well on tasks that are severely ill-conditioned, such as recurrent networks, deep autoencoders, physics-informed neural networks, and a 136 million–parameter physics-informed diffusion model, often achieving lower training losses than established baselines including Adam, Muon, and SOAP. The authors also provide a Python implementation.