Skip to content

L2CO Optimizers¤

Bare optimizers compatible with the L2CO library.

l2co-optimizers is the optimizer half of the L2CO ecosystem, as l2co-tasks is the task half. It provides: - a name registry of ready-to-run optimizers: the optax gradient methods, the evosax distribution- and population-based algorithms, the optimistix and scipy minimisers, IPOPT, plus SHADE, TuRBO, an RBF trust region, per-evaluation-key L-BFGS and random search; - the UpdateClass container every registry factory returns; - the OptimizationStep spec that names an optimizer with its hyperparameters and stopping criteria; - ready-made Hydra optimizer configs.

It also ships what each optimizer exposes to a loop that switches between optimizers: its state-transfer ports, and optimizer_parts, which unpacks it into an optax transform or an (init, ask, tell) triple. The switching layer itself (the SubOpt adapter, the handshake policy, menu-dispatched loss evaluation), the meta-optimization strategies (l2co, rl2co, agentic-l2co) and the bridge to tasks live in l2co.

Statement of need¤

Learning-to-optimize and optimizer-selection research needs many optimizers behind one calling convention, so that a selector can switch between them mid-run. l2co-optimizers provides that convention without the meta-learning stack: - every optimizer is built by optimizer_mapping(name)(model=..., loss_fn=..., pass_rng=..., opt_hash=..., bounded=..., stop_fn=...); - every one runs a budget through the same UpdateClass.run and reports into the same HistoryState; - every one except the scipy minimisers and IPOPT also steps through the same (params, opt_state, key) carry, reporting each step as an OptHistory. Those seven own their loop and run the whole budget in one call (ADRs 0002 and 0003).

It depends on neither l2co nor l2co-tasks, and has no notion of a task: factories take model, loss_fn and pass_rng as keywords. l2co is the bridge that unpacks an l2co_tasks.Task into them (l2co ADR 0018).

Available optimizers¤

Every optimizer below is built by name through optimizer_mapping(name). Names are normalized (non-alphanumerics stripped, lowercased), so "rbf_trust_region" and "rbftrustregion" resolve to the same entry. Meta-optimizers (l2co, rl2co) are not built in: they register themselves when their package is imported.

Name Algorithm Family Backend
adabelief AdaBelief Gradient optax
adadelta AdaDelta Gradient optax
adafactor Adafactor Gradient optax
adagrad AdaGrad Gradient optax
adam Adam Gradient optax
adamax AdaMax Gradient optax
adamaxw AdaMax with decoupled weight decay Gradient optax
adamw AdamW Gradient optax
adan Adan Gradient optax
amsgrad AMSGrad Gradient optax
fromage Fromage Gradient optax
lamb LAMB Gradient optax
lars LARS Gradient optax
lion Lion Gradient optax
nadam NAdam (Adam with Nesterov momentum) Gradient optax
nadamw NAdamW (AdamW with Nesterov momentum) Gradient optax
noisysgd Noisy SGD Gradient optax
novograd NovoGrad Gradient optax
optimisticadam Optimistic Adam Gradient optax
optimisticgradientdescent Optimistic gradient descent Gradient optax
radam RAdam Gradient optax
rmsprop RMSProp Gradient optax
rprop Rprop Gradient optax
sgd SGD Gradient optax
signsgd signSGD Gradient optax
sm3 SM3 Gradient optax
yogi Yogi Gradient optax
lbfgs L-BFGS, with a fresh PRNG key per linesearch evaluation on stochastic objectives Quasi-Newton optax + built-in
bfgs BFGS with a backtracking Armijo line search Quasi-Newton optimistix
dfp DFP with a backtracking Armijo line search Quasi-Newton optimistix
nonlinearcg Nonlinear conjugate gradient (Polak-Ribiere by default; Fletcher-Reeves, Hestenes-Stiefel, Dai-Yuan) with a backtracking Armijo line search Gradient optimistix
tnc Truncated Newton (TNC): a line-search Newton method on finite-difference Hessian-vector products Newton-type scipy
trustkrylov Newton trust region with a Krylov (GLTR) subproblem solver; Hessian-vector products by finite differences of gradients Newton-type scipy
slsqp Sequential least-squares quadratic programming (SLSQP): SQP with a dense BFGS Hessian and an L1 merit line search Quasi-Newton scipy
trustconstr trust-constr: trust-region SQP (an interior-point method when there is a box) with a dense BFGS Hessian Quasi-Newton scipy
ipopt IPOPT: primal-dual interior point with a filter line search and a limited-memory quasi-Newton Hessian (Wächter & Biegler 2006) Quasi-Newton IPOPT, through casadi
ars Augmented Random Search Distribution-based evosax
asebo ASEBO Distribution-based evosax
cmaes CMA-ES Distribution-based evosax
crfmnes CR-FM-NES Distribution-based evosax
des Discovered ES Distribution-based evosax
esmc ESMC Distribution-based evosax
gradientlessdescent Gradientless Descent Distribution-based evosax
guidedes Guided ES Distribution-based evosax
hillclimbing Hill climbing Distribution-based evosax
iamalgamfull iAMaLGaM (full covariance) Distribution-based evosax
iamalgamunivariate iAMaLGaM (univariate) Distribution-based evosax
lmmaes LM-MA-ES Distribution-based evosax
maes MA-ES Distribution-based evosax
noisereusees Noise-Reuse ES Distribution-based evosax
openes OpenAI-ES Distribution-based evosax
persistentes Persistent ES Distribution-based evosax
pgpe PGPE Distribution-based evosax
rmes Rm-ES Distribution-based evosax
sepcmaes Sep-CMA-ES Distribution-based evosax
simplees Simple ES Distribution-based evosax
simulatedannealing Simulated annealing Distribution-based evosax
snes SNES Distribution-based evosax
xnes xNES Distribution-based evosax
differentialevolution Differential Evolution Population-based evosax
diffusionevolution Diffusion Evolution Population-based evosax
gesmrga GESMR-GA Population-based evosax
mr15ga MR15-GA Population-based evosax
pso Particle Swarm Optimization Population-based evosax
samrga SAMR-GA Population-based evosax
simplega Simple GA Population-based evosax
neldermead Nelder-Mead downhill simplex; the population is the simplex (dimensionality + 1 vertices) Population-based optimistix
shade SHADE, with optional turning-based mutation (Tanabe & Fukunaga 2013; Sun et al. 2020) Population-based built-in (evosax API)
turbo TuRBO trust-region Bayesian optimization (Eriksson et al. 2019) Model-based built-in
rbf_trust_region RBF-surrogate trust-region search (ORBIT / DYCORS family) Model-based built-in
cobyqa COBYQA: derivative-free trust region on quadratic interpolation models (Ragonneau & Zhang) Model-based scipy
powell Powell's conjugate direction method (derivative-free line searches) Direct search scipy
randomsearch One-shot random search Random built-in

The four optimistix entries (bfgs, dfp, nonlinearcg, neldermead) run as plain registry entries only: they evaluate the objective themselves, so optimizer_parts cannot unpack them into a switching menu. They bill the evaluations optimistix actually makes (one per step for the gradient solvers; for Nelder-Mead n + 1 on the first step, 2 per step and n + 3 on a shrink), never stop early on their own convergence test, and clip into bounded before each evaluation. See ADR 0001.

The six scipy entries (cobyqa, powell, tnc, trustkrylov, slsqp, trustconstr) are plain registry entries too, for a different reason: scipy owns the optimization loop, so each run hands its whole budget to scipy.optimize.minimize inside one host callback. One iteration is one evaluation (value, or value and gradient), billed one. When scipy finishes before the budget, every remaining iteration re-evaluates its final point; if scipy fails, the run stays at the best point found. A stop_fn raises. See ADR 0002, and ADR 0003 for SLSQP and trust-constr.

ipopt runs the same way, on the same driver, through casadi, whose wheels bundle IPOPT. IPOPT uses its own limited-memory quasi-Newton Hessian, and a value request and a gradient request at the same point are one evaluation. See ADR 0003.

Three of these get expensive in high dimensions. COBYQA's cost per evaluation grows steeply with the dimensionality, and SLSQP's and trust-constr's do from about a thousand dimensions, so high-dimensional runs of these three can exceed a cluster's wall-clock limit.

To add your own optimizer to the registry, see Register your own optimizer.