Every few years, a derivative-free neural network optimization algorithm gets some hype. I'd bet my life savings that none of them ever make an impact.
Derivative-free optimization can be useful for genuinely discontinuous objectives [1], but common neural network objectives are smooth and/or Lipschitz.
The gradient is useful. Instead of trying random directions and hoping that one of them is an improvement, it tells you where to go. The more parameters you have, the more useful it becomes.
A strict complexity gap between gradient-based and derivative-free Lipschitz convex optimization has been suspected for decades and recently proved (using AI, [2]). Neural net optimization is nonconvex, but not radically different.
IMO, a more promising direction is gradient-based optimizers specialized to the neural network structure, like Muon [3].
I'm skeptical as to whether zeroth-order methods really lend themselves to a Bitter Lesson argument. First-order methods don't explore the loss landscape optimally, but the loss function tends to be nonconvex, and zeroth-order methods don't address that issue head on. Dust smooths, and so do applicable first-order methods. Remove the nonconvexity issue, and I suspect Dust's purported advantages evaporate (based on published theoretical work), so it's pretty odd to me that the paper never discusses convexity.
I could buy that this method scales better than previous zeroth-order methods, and that's interesting, but it doesn't seem like enough of a moat to keep improved first-order methods from drinking its milkshake, except in cases where a zeroth-order method is already a primary option: the network needs to call a simulator that doesn't expose gradient-like information. (In cases where gradients don't exist, I'd still argue for other options, e.g., Clarke-generalized gradients where applicable, so long as those can be computed with the available information. I know this technology has been published for automatic differentiation, so I would imagine it could be incorporated into backprop and used with a suitable optimization algorithm.)
show comments
mkaic
I'm excited to see a new thing in the Zero-Order Optimization (ZOO) world. I see most ZOO methods not as a "replacement for backprop" the way they are often marketed, but as a technique that can potentially work well in regimes where backprop is fundamentally weak. One of my favorite papers [0] in recent memory, for instance, is about using central-difference random gradient estimation (CD-RGE) to train large RNNs without using backprop-through-time. They also show it works decently for hard-to-optimize differentiable-neural-computers. Since reading that paper, I've had a lot of fun trying to apply this technique to settings I would describe as "a traditional differentiable neural network being applied in a non-traditional environment where you can't just call loss.backward and hope for the best". I am excited to try Dust out on my personal project now :)
usernametaken29
> There are many interesting open questions. The first is whether, and how, Dust can find better directions than backprop’s first-order gradient
Both algorithms are bound by the same Pareto frontier based on the Empirical Risk Minimisation Principle, so they’re already on the same trajectory.
Interestingly backprop is limited by conditioning of the Hessian matrix in order to converge (differentiate correctly). So removing this limitation is actually a great step.
I’m excited to see a comeback of evolutionary methods because they’re much more general, albeit costly and naive.
We’re now very close to what can be described best as brute forcing the Pareto frontier out of our datasets. Not sure that’s what we want but I have no better ideas either.
show comments
polyomino
Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains?
It would be cool to apply this to different stages and see if that affects the learning trajectory
ghm2180
Intereting. There was a completely different take on how to "do" back propagation using optics at YC papers video the other day(I can't find the video). The technique presented diffusion networks with optics. If a hardware technique to make computation more efficient emerges, a lot of methods like this could apply no?
soltanov
I would like to see wall clock time, energy, peak memory and downstream quality compared at equal loss. Until the effiency gap closes, this is an interesting research direction.
sim04ful
For me credit assignment without a global clock is extremely important, i believe it's a necessary step towards end to end training of models In-materio.
api
It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?
show comments
wg0
It's computationally expensive and infeasible but what's the upside?
Genuine question due to unfamiliarity with the subject.
show comments
ammarmalik17
Bigger Dust models needing fewer population samples?
oofbey
This is super yawn-worthy. Instead of backprop for the exact gradient you can run forward passes a thousand times with perturbed weights and get a Monte Carlo estimate of the gradient. Not very clever. Extremely NOT useful.
But I guess the industry is littered with techniques for computing the same thing but vastly slower that some people find interesting. Homomorphic encryption. Zero knowledge proofs. Blockchain computing. Except in those cases there might be a legitimate reason to use it occasionally.
ContinuityLab
Moving beyond traditional backpropagation for transformer pretraining opens up fascinating avenues for alternative learning dynamics and architectural efficiency.
vivzkestrel
- stupid question from a neural network rookie
- isnt the whole point of back propoagation so that you dont guess weights by brute forcing them since that is computationally infeasible once you go beyond a dozen weights?
- if you dont use backpropogation, how exactly are the initial weights assigned them if they are not random values?
AIorNot
Hmm I was more interested in this alternative to backprop:
Every few years, a derivative-free neural network optimization algorithm gets some hype. I'd bet my life savings that none of them ever make an impact.
Derivative-free optimization can be useful for genuinely discontinuous objectives [1], but common neural network objectives are smooth and/or Lipschitz.
The gradient is useful. Instead of trying random directions and hoping that one of them is an improvement, it tells you where to go. The more parameters you have, the more useful it becomes.
A strict complexity gap between gradient-based and derivative-free Lipschitz convex optimization has been suspected for decades and recently proved (using AI, [2]). Neural net optimization is nonconvex, but not radically different.
IMO, a more promising direction is gradient-based optimizers specialized to the neural network structure, like Muon [3].
[1] https://arxiv.org/abs/2202.00817
[2] https://arxiv.org/abs/2607.13335
[3] https://jeremybernste.in/writing/deriving-muon
I'm skeptical as to whether zeroth-order methods really lend themselves to a Bitter Lesson argument. First-order methods don't explore the loss landscape optimally, but the loss function tends to be nonconvex, and zeroth-order methods don't address that issue head on. Dust smooths, and so do applicable first-order methods. Remove the nonconvexity issue, and I suspect Dust's purported advantages evaporate (based on published theoretical work), so it's pretty odd to me that the paper never discusses convexity.
I could buy that this method scales better than previous zeroth-order methods, and that's interesting, but it doesn't seem like enough of a moat to keep improved first-order methods from drinking its milkshake, except in cases where a zeroth-order method is already a primary option: the network needs to call a simulator that doesn't expose gradient-like information. (In cases where gradients don't exist, I'd still argue for other options, e.g., Clarke-generalized gradients where applicable, so long as those can be computed with the available information. I know this technology has been published for automatic differentiation, so I would imagine it could be incorporated into backprop and used with a suitable optimization algorithm.)
I'm excited to see a new thing in the Zero-Order Optimization (ZOO) world. I see most ZOO methods not as a "replacement for backprop" the way they are often marketed, but as a technique that can potentially work well in regimes where backprop is fundamentally weak. One of my favorite papers [0] in recent memory, for instance, is about using central-difference random gradient estimation (CD-RGE) to train large RNNs without using backprop-through-time. They also show it works decently for hard-to-optimize differentiable-neural-computers. Since reading that paper, I've had a lot of fun trying to apply this technique to settings I would describe as "a traditional differentiable neural network being applied in a non-traditional environment where you can't just call loss.backward and hope for the best". I am excited to try Dust out on my personal project now :)
> There are many interesting open questions. The first is whether, and how, Dust can find better directions than backprop’s first-order gradient
Both algorithms are bound by the same Pareto frontier based on the Empirical Risk Minimisation Principle, so they’re already on the same trajectory. Interestingly backprop is limited by conditioning of the Hessian matrix in order to converge (differentiate correctly). So removing this limitation is actually a great step. I’m excited to see a comeback of evolutionary methods because they’re much more general, albeit costly and naive. We’re now very close to what can be described best as brute forcing the Pareto frontier out of our datasets. Not sure that’s what we want but I have no better ideas either.
Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains? It would be cool to apply this to different stages and see if that affects the learning trajectory
Intereting. There was a completely different take on how to "do" back propagation using optics at YC papers video the other day(I can't find the video). The technique presented diffusion networks with optics. If a hardware technique to make computation more efficient emerges, a lot of methods like this could apply no?
I would like to see wall clock time, energy, peak memory and downstream quality compared at equal loss. Until the effiency gap closes, this is an interesting research direction.
For me credit assignment without a global clock is extremely important, i believe it's a necessary step towards end to end training of models In-materio.
It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?
It's computationally expensive and infeasible but what's the upside?
Genuine question due to unfamiliarity with the subject.
Bigger Dust models needing fewer population samples?
This is super yawn-worthy. Instead of backprop for the exact gradient you can run forward passes a thousand times with perturbed weights and get a Monte Carlo estimate of the gradient. Not very clever. Extremely NOT useful.
But I guess the industry is littered with techniques for computing the same thing but vastly slower that some people find interesting. Homomorphic encryption. Zero knowledge proofs. Blockchain computing. Except in those cases there might be a legitimate reason to use it occasionally.
Moving beyond traditional backpropagation for transformer pretraining opens up fascinating avenues for alternative learning dynamics and architectural efficiency.
- stupid question from a neural network rookie
- isnt the whole point of back propoagation so that you dont guess weights by brute forcing them since that is computationally infeasible once you go beyond a dozen weights?
- if you dont use backpropogation, how exactly are the initial weights assigned them if they are not random values?
Hmm I was more interested in this alternative to backprop:
https://news.ycombinator.com/item?id=49701182
Definitely more than meets the eye.
Remember kids, whoever is pitching space searching via complete enumeration just wants your wallet.
Dust's 243M model beating a 120x smaller one at most population sizes is the surprising part; bigger nets got more population-efficient, not less.