18 comments

  • blt 18 hours ago

    Every few years, a derivative-free neural network optimization algorithm gets some hype. I'd bet my life savings that none of them ever make an impact.

    Derivative-free optimization can be useful for genuinely discontinuous objectives [1], but common neural network objectives are smooth and/or Lipschitz.

    The gradient is useful. Instead of trying random directions and hoping that one of them is an improvement, it tells you where to go. The more parameters you have, the more useful it becomes.

    A strict complexity gap between gradient-based and derivative-free Lipschitz convex optimization has been suspected for decades and recently proved (using AI, [2]). Neural net optimization is nonconvex, but not radically different.

    IMO, a more promising direction is gradient-based optimizers specialized to the neural network structure, like Muon [3].

    [1] https://arxiv.org/abs/2202.00817

    [2] https://arxiv.org/abs/2607.13335

    [3] https://jeremybernste.in/writing/deriving-muon

    • accurrent 9 hours ago

      It's still important to support research in this direction cause us meatbags cost a LOT less energy to train than GPTs even if you assume that it takes 30 years to train a PhD. Our current training methods honestly leave a lot to be desired. We almost certainly dont do full backpropagation.

      • mkehrt 5 hours ago

        I’m not sure that’s true. We have literal billions of years of training.

        • howunfortunate 2 hours ago

          Human "pre-training" is vastly underappreciated.

          Hell, animals are an even better example. Many animals pop out of the womb and start walking and eating and acting just like an adult!

          • sva_ 1 hour ago

            One could argue that humans are more adaptable to novel environments

          • culi 1 hour ago

            If you wanna get philosophical then the decades of research that humans did is also AI "pre-training". And all of that is built on knowledge that humans have spent our whole existence amassing so maybe all of human "pre-training" is also AI pre-training.

            A fun but silly exercise.

          • ashf023 2 hours ago

            I'm sure we don't literally do backpropagation, but differential equations that settle toward stationary states show up a lot in biology

            • throw310822 2 hours ago

              You need to factor in the cost of training all those PhDs that never end up producing much of interest though you don't get to pick the best afterwards and claim all it took was to train him/her. Plus, once trained, it's productive only for 6 subjective hours per day (including weekends and holidays). Might have costed gazillions to train a frontier LLM but it works for millions of hours/day.

              • empath75 7 hours ago

                PHD inference doesn't scale, though.

                • refulgentis 7 hours ago

                  I think it's fun back-of-the-napkin sometimes to compare meat to matmuls but ultimately fallacious to its core, making it an intellectual tarpit.

                  It's much harder to argue with the math and empirical results.

                  • abecedarius 6 hours ago

                    Energy cost may be a pointless comparison since we don't eat electricity, but it's very reasonable to look at sample efficiency and ask what we have there that our theories miss.

                • syntacticsalt 17 hours ago

                  Even for non-smooth or discontinuous objectives, I'd still reach for methods that use gradient-like information over zeroth-order methods. For non-smooth objectives, Clarke-generalized subdifferentials have been pretty effective outside of ML, and have been used in automatic differentiation contexts at least 10-15 years ago. A carelessly quick literature search suggested conservative gradients, too.

                  For discontinuous objectives, I know there's been work on using envelope approximations, but the little I'm aware of in that work was in low-dimensional settings where the structure of the discontinuity was known explicitly. On the other extreme, lack of continuity comes up all the time in infinite-dimensional, PDE-constrained optimization, and some methods rely on tangent cones or various generalized notions of subdifferentiability (e.g., Mordukhovich, Bouligand) to demonstrate convergence. Admittedly, that work was somewhat outside my area of expertise, so I may be getting the details there slightly wrong, but the broad point stands that even in those settings, some directional information can be obtained and used profitably without resorting to zeroth-order methods.

                  • armcat 11 hours ago

                    Kind of like random projections. Pre ChatGPT there would every now and then come a paper that states something along the lines of: instantiate a large randomly populated matrix, multiply by the input, and win! It kinda makes sense because you "stretch out" the space and separation boundaries become easier, but I have never seen it fully utilised in production in any meaningful way. Anyone remember the DANs (Deep Averaging Networks)?

                    • numeri 2 hours ago

                      This is used quite often in SoTA quantization algorithms, which definitely count as production

                    • rtmkrptn 1 hour ago

                      It is like Monte Carlo integration and the curse of dimensionality. Yes, you can use this for basically anything but it pays off only if the object you are trying to integrate is multidimensional, otherwise traditional methods outperforms

                      • torginus 2 hours ago

                        Isn't the 'we need gradients' a foregone conclusion?

                        Gradients can be calculated numerically, meaning that any method that samples the cost function and makes optimization decisions based on that can actually compute gradients if it needs to.

                        • vatsachak 18 hours ago

                          Random selection seems to favor post training

                          https://arxiv.org/pdf/2603.12228

                          • machiaweliczny 9 hours ago

                            If it's cheaper then I guess it will might get practical. Also more likely that mind uses simpler techniques more similar to these.

                            I wonder though if someone tested mutating training objective though, like keeping original loss/goal and somehow defining loss differently and then comparing against original. This intuitively feels like how mind tries to handle difficult tasks.

                            • okintheory 10 hours ago

                              "A strict complexity gap between gradient-based and derivative-free Lipschitz convex optimization" has been known for decades. It's covered in standard textbooks like Nesterov's "Introductory lectures on convex optimization". That AI result is about establishing the gap in a fairly niche accuracy regime.

                              • blt 5 hours ago

                                Thanks for the correction, I had a feeling that should be true but the headline was at top of search results and I'm not close enough to the field to know off the top of my head.

                              • cubefox 12 hours ago

                                > Derivative-free optimization can be useful for genuinely discontinuous objectives [1], but common neural network objectives are smooth and/or Lipschitz.

                                There is even an analog to the continuous derivative for discrete binary functions, called "Boolean variation": https://proceedings.neurips.cc/paper_files/paper/2024/hash/7...

                                Like for derivatives, there is a chain rule for Boolean variations, so you can use something like backpropagation, but without needing any expensive floating point math. Though I don't think this has been used much so far. There must be some other downside.

                                • p1esk 13 hours ago

                                  And yet the best learning algorithm we have is derivative-free!

                                  • ghm2180 11 hours ago

                                    Yes, SVMs and Random Forests still rule the many worlds of classification problems often not in the spotlight.

                                    • cyber_kinetist 10 hours ago

                                      Support Vector Machines involve solving a large QP optimization problem... which is often done by running a variant of gradient descent (or one of the second- or quasi-second-order optimization algorithms, which involves finding the Hessian as well as the gradient)

                                      • Maxatar 2 hours ago

                                        I was under the impression that solving QP optimization problems (subject to constraints) was largely performed by SMO.

                                      • lynndotpy 9 hours ago

                                        Back when I was a researcher, I had a classification problem where the default random forest classifier built into scikitlearn completely dominated the neural method. The best numbers we got were something like a ~30% improvement from baseline for the neural network to ~95% for the random forest.

                                        The disparity was so large that I was certain I must have made a mistake and I spent a few hours debugging, and then a few hours more trying different neural architectures.

                                        Turns out this is just a super common experience for anyone in NNs who would also try the more established learning algorithms.

                                      • cubefox 12 hours ago

                                        Perhaps the optimizer is worse but the architecture or objective function is better.

                                        • WithinReason 12 hours ago

                                          Imagine how good it would be with derivatives

                                          • YeGoblynQueenne 10 hours ago

                                            Bayes Optimal Classifier?

                                        • syntacticsalt 17 hours ago

                                          I'm skeptical as to whether zeroth-order methods really lend themselves to a Bitter Lesson argument. First-order methods don't explore the loss landscape optimally, but the loss function tends to be nonconvex, and zeroth-order methods don't address that issue head on. Dust smooths, and so do applicable first-order methods. Remove the nonconvexity issue, and I suspect Dust's purported advantages evaporate (based on published theoretical work), so it's pretty odd to me that the paper never discusses convexity.

                                          I could buy that this method scales better than previous zeroth-order methods, and that's interesting, but it doesn't seem like enough of a moat to keep improved first-order methods from drinking its milkshake, except in cases where a zeroth-order method is already a primary option: the network needs to call a simulator that doesn't expose gradient-like information. (In cases where gradients don't exist, I'd still argue for other options, e.g., Clarke-generalized gradients where applicable, so long as those can be computed with the available information. I know this technology has been published for automatic differentiation, so I would imagine it could be incorporated into backprop and used with a suitable optimization algorithm.)

                                          • Buttons840 11 hours ago

                                            Is it fair to define Reinforcement Learning (RL) as forcing a gradient onto a system / simulator?

                                            Though, I suppose RL has non-gradient based methods too.

                                            • Den_VR 16 hours ago

                                              Orders-of-magnitude improvements in compute efficiency are needed to become a practical replacement for backprop… but those improvements are coming.

                                              As a mixture, could activation-space search produce useful teaching targets for backprop?

                                              Zeroth-order search would discover candidates, first-order learning would consolidate them. The potentially valuable step is converting a sparse judgment into a reusable training target.

                                              This also changes the relevance of convexity.

                                              • syntacticsalt 16 hours ago

                                                I'm not sure I follow why your argument changes the relevance of convexity? If the problems were convex, I don't think we'd be having this discussion -- first-order methods would tend to win.

                                            • mkaic 7 hours ago

                                              I'm excited to see a new thing in the Zero-Order Optimization (ZOO) world. I see most ZOO methods not as a "replacement for backprop" the way they are often marketed, but as a technique that can potentially work well in regimes where backprop is fundamentally weak. One of my favorite papers [0] in recent memory, for instance, is about using central-difference random gradient estimation (CD-RGE) to train large RNNs without using backprop-through-time. They also show it works decently for hard-to-optimize differentiable-neural-computers. Since reading that paper, I've had a lot of fun trying to apply this technique to settings I would describe as "a traditional differentiable neural network being applied in a non-traditional environment where you can't just call loss.backward and hope for the best". I am excited to try Dust out on my personal project now :)

                                              • usernametaken29 19 hours ago

                                                > There are many interesting open questions. The first is whether, and how, Dust can find better directions than backprop’s first-order gradient

                                                Both algorithms are bound by the same Pareto frontier based on the Empirical Risk Minimisation Principle, so they’re already on the same trajectory. Interestingly backprop is limited by conditioning of the Hessian matrix in order to converge (differentiate correctly). So removing this limitation is actually a great step. I’m excited to see a comeback of evolutionary methods because they’re much more general, albeit costly and naive. We’re now very close to what can be described best as brute forcing the Pareto frontier out of our datasets. Not sure that’s what we want but I have no better ideas either.

                                                • Llamamoe 1 hour ago

                                                  > Interestingly backprop is limited by conditioning of the Hessian matrix in order to converge (differentiate correctly)

                                                  What does this mean?

                                                  • stevemk14ebr 19 hours ago

                                                    Pareto frontier the Pareto frontier and then with a little Pareto frontier you might get to the Pareto frontier

                                                    • usernametaken29 18 hours ago

                                                      P^2

                                                      • AIorNot 18 hours ago

                                                        Its just an insufferable AI way to say best set of options given

                                                        • tomrod 18 hours ago

                                                          Waaay older than AI. Basic econ 101,it is.

                                                          • girvo 13 hours ago

                                                            Though it being used constantly in AI-related topics from ~2024 onwards is very much a function of LLM output, I would argue.

                                                            • tomrod 10 hours ago

                                                              Oh for sure. As has em-dashes -- a glorious tidbit primarily used by writers in the past to circumvent parenthetical clauses. Also the incredibly abused "load-bearing."

                                                              Meanwhile, I just want people to think about the production possibilities frontier.

                                                          • usernametaken29 18 hours ago

                                                            Not really. The Pareto frontier is itself a distribution over all optimisation paths and dominance is a real and useful property we test when researching evolutionary methods. It’s not chosen to be insufferable but to try to be precise.

                                                    • polyomino 1 day ago

                                                      Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains? It would be cool to apply this to different stages and see if that affects the learning trajectory

                                                      • ghm2180 11 hours ago

                                                        Intereting. There was a completely different take on how to "do" back propagation using optics at YC papers video the other day(I can't find the video). The technique presented diffusion networks with optics. If a hardware technique to make computation more efficient emerges, a lot of methods like this could apply no?

                                                        • soltanov 14 hours ago

                                                          I would like to see wall clock time, energy, peak memory and downstream quality compared at equal loss. Until the effiency gap closes, this is an interesting research direction.

                                                          • sim04ful 15 hours ago

                                                            For me credit assignment without a global clock is extremely important, i believe it's a necessary step towards end to end training of models In-materio.

                                                            • api 1 day ago

                                                              It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?

                                                              • vatsachak 1 day ago

                                                                Not necessarily, backprop is highly parallelizable since it is just a bunch of matrix mults.

                                                                Something like Dust skips the backward pass on backprop. But other techniques like Neural Predictive Coding can be completely asynchronous, each "weight" can fire independent of those far away from it. Innocenti, et. al have shown that NPC gradients converge to backprop within a certain "regime".

                                                                The win with asynchronous techniques like NPC is that you do not need the extreme co-ordination that backprop requires and hence should be computationally much easier given the right device.

                                                                Although at this point the industry has so much money in the forward-backward pass system that I doubt a backprop successor would win unless someone makes NPC hardware feasible and can prove scaling up to billions of params

                                                                • tyromaniac 22 hours ago

                                                                  Also has the "advantage" of being slightly more biologically plausible as the optimization happens locally rather than globally.

                                                                  That idea was taken further by N'dri et al in PCL, in which "activation energy" was minimized as well, and inhibitory neurons added https://www.nature.com/articles/s41467-025-64234-z.pdf

                                                                  While trying to find the link for that I stumbled upon

                                                                  https://arxiv.org/pdf/2605.12732

                                                                  Which also looks pretty interesting

                                                                  • ACCount39 11 hours ago

                                                                    I think this is the main thing I want out of "non-backprop ML" research - figuring out how the whole class of algorithms behaves. And then applying that to figure out how the brain implements its own deep learning.

                                                                    If we can figure out how the brain's learning dynamics function well enough? We could figure out how to interface with them and extend them.

                                                                  • ACCount39 23 hours ago

                                                                    The reason why I don't see the promise for ML-only applications is that the coordination backprop requires comes very cheap to us.

                                                                    "Much easier given the right device" - the "right" there just isn't shaped like the devices we actually build. And the price of "not having backprop" is usually expending more FLOPs, getting worse sample efficiency, etc.

                                                                    The biggest "device" that doesn't do backprop is the brain, and that's because the brain doesn't have the connectivity or the coordination to pull it off. Both of those are "expensive" for something like it to implement. Cheap for us though. We aren't stuck with neurons that only get locally available information and have to implement learning rules based on that. So, skill issue?

                                                                    • ummonk 4 hours ago

                                                                      What about applications like distributed computing using volunteer computers, instead of datacenters? You can't really do that with normal backprop approaches. E.g. for LC0 the training data is generated in a distributed manner, but the neural nets themselves are trained centrally.

                                                                      • vatsachak 22 hours ago

                                                                        I mean co-ordination requires energy though. The brain wattage looks at GPUs and says "skill issue". But you're right that we look at natural energy production techniques and say "skill issue"

                                                                        • dharma1 12 hours ago

                                                                          I think both can be true - ie. We can harness energy from the sun at much bigger scale than nature has so far, yet we can still learn a lot from how nature has spent massive parallel search over millions of years on coming up with incredible nano engineering we still have no clue how to do

                                                                          • ACCount39 21 hours ago

                                                                            Even today's LLMs suddenly get power-competitive when you compare by "power per task". Sure, a GPU can draw 1000W under load. But it also works very fast, and doesn't have to spend any time on things like "sleep".

                                                                            The trick about comparing the two is that different things are expensive to different substrates.

                                                                            Coordination is cheap for GPGPU and expensive for brain. When you have a fixed number of reusable general purpose computational units, coordinating execution is more natural than not coordinating execution, and the power cost is nil. When your computational units are independent, purpose specific, and fully embedded into the data path, coordinating them can get less natural and, frankly, optional. When wiring is expensive, coordination can become expensive in turn.

                                                                            Another thing in the same "cheap for GPGPU but expensive for brain" regime is bandwidth. Look no further than optic nerve to see just how hard it is for nerves to push any appreciable amount of data. Another thing is connectivity. For GPGPU, global connectivity is natural - but the brain has to pay in physical wires for all the connectivity it has, and, see "bandwidth": it doesn't have any good wires. Yet another thing is weight reuse: a big part of why humans get "handedness" is that the brain can't just reuse the motion control circuitry for one hand for another nearly identical hand.

                                                                            And the final thing I can name off the top of my head is memory - but specifically, memory capable of fast R/W. The capacity of human "working memory" is a disgrace, and not because there was no use for more. Humans rapidly lose visual fidelity of representations for objects they aren't directly looking at, and not at all because "being able to check how things looked 2 seconds ago" is useless. Those capabilities were just too expensive for the substrate to afford them easily.

                                                                            It's why brains, broadly, favor dataflow-like and SSM-like dynamics, with largely fixed asynchronous dataflows and recurrence over updated local information - instead of something that would require a lot of global connectivity and transformer-like many-to-many attention ops. SSM is not necessarily the "best" tool for the job in ML land, for most jobs - but when you struggle to fit "attention" into your connectivity/bandwidth budget, and your memory is extremely expensive but hard-coupled to processing, SSM starts looking very appealing.

                                                                            Now, something that might be expensive for GPGPU but cheap for brain, for once? Online learning. Maybe it's substrate dependent, or maybe it's going to get cheap in GPU land too once we figure out the trick. But so far? No one figured out how to make it cheap, stable and usable. You'd be lucky to get "pick one".

                                                                            • vatsachak 21 hours ago

                                                                              I mean the reason why models are using less energy is because they are getting smarter per token and also engineering algorithms/chips that make inference cheaper.

                                                                              If we could have success with spiking neural networks in silico they would take even less energy, because they don't require global co-ordination. Co-ordination is information and "information = energy by the second law of thermodynamics" is my crank proof

                                                                              Also the brain has way more parameters than LLMs and also has different neurotransmitters, loops, branching etc so they probably have WAY more capacity than LLMs.

                                                                              But coding output/W LLMs have us beat

                                                                              • ACCount39 14 hours ago

                                                                                Coordination is fuck all bits worth of control information broadcast widely. It's very cheap to us.

                                                                                I frankly don't believe in spiking neural networks giving any advantages over what we have. It's a different way to implement ANNs, but "different" isn't "better". It's how the brain does things, sure, but the answer to "why the brain does what it does" is "workarounds for being made of flesh issues" at least half the time.

                                                                                I can believe in brain having more capacity than frontier LLMs quite easily. We know a single BNN neuron can have the expressiveness of many ANN neurons. And well leveraged overparametrization + compute overhang could explain a decent chunk of the apparent sample efficiency edge.

                                                                                But that apparent "extra capacity" could also be tied up in things like neurons having to contend with metabolism, in brain's learning algorithms being noisy, in brain having to use neuron circuits to implement "hot memory", etc - instead of contributing only to performance.

                                                                        • janalsncm 22 hours ago

                                                                          Knowing nothing about this, I wonder if it could be useful in situations where we can’t reliably sync with all the workers. Something like folding@home, where all the workers are just shaking weights and if one of them finds a winner it uploads to the central server?

                                                                          • vatsachak 22 hours ago

                                                                            No real advantage over Neural Nets here; backprop matmuls can be calculated layer by layer so you can chunk backprop across different machines. The real advantage comes from energy savings, you require no global co-ordination

                                                                            • janalsncm 22 hours ago

                                                                              Imagine we did that, split up a model layers as A->B->C. C will need to wait for B to compute a forward pass, which is waiting for A to compute its forward pass. To compute the forward pass, B needs all of the outputs from A, which is an upload and a download (maybe these can be done concurrently).

                                                                              Then A waits for B to compute its backwards pass, which is waiting for C to do the same thing. Again you are sending around potentially gigabytes of data.

                                                                              This is in contrast to mining bitcoins for example which doesn’t require any coordination from miners because their work is completely independent, and the answer is very small compared to the work needed to get it.

                                                                              • vatsachak 21 hours ago

                                                                                Yep. You need to transport all the weights at the boundary regardless of Backprop/NPC.

                                                                                But the cool thing is that if your NN is split into mostly self contained chunks then you can go widthwise parallel.

                                                                                An architecture like MOE exploits this fact so that the active weights during pre-training you're backproping only through active experts

                                                                              • spindump8930 20 hours ago

                                                                                The problem (and contrast with other approaches) is that mat muls requires synchronization. Arranging your networking and training structure to maximize compute and minimize communication is the main craft of ML training infra folks. In your example, yes you can compute layers on different machines (i.e. Tensor Parallelism), but you must be very careful in how you arrange it.

                                                                            • dharma1 12 hours ago

                                                                              I think Jeff Dean is right in that we will see much more specialised silicon in the future.

                                                                              If something more bio inspired ie. predictive coding and in-memory compute fundamentally makes continual learning and much lower energy consumption possible there will be specialised hardware for it at some point

                                                                              FWIW I think the brain has multiple “learning rules” and operates at multiple timescales

                                                                              • strbean 22 hours ago

                                                                                Would these alternatives to backprop make it more feasible to have constant live-training going on in a model? Giving it something akin to neuro-plasticity?

                                                                                • nbutton762 21 hours ago

                                                                                  One aspect of how current training and continual learning are somewhat at odds is that the memory required to train a model is often times 2-3x the memory required to just run it (probably not as bad for PEFT, not sure).

                                                                                  DUST does have an advantage specifically along those lines because it doesn't have to save a ton of intermediate state other than each layer's input activations during a single forward pass.

                                                                                  There are many other issues that this algorithm does not address thoigh like catastrophic forgetting. it's still operating on a transformer which contains no inherent mechanism for selecting the relative value of a training step based on current knowledge, nor does it have segmentation of functionalities with specialized areas used for specific things that can be sequestered off and ignore new updates (we do not risk forgetting how to walk as we increase our French vocabulary)

                                                                                  • vatsachak 21 hours ago

                                                                                    Models suffer from "catastrophic forgetting" if you train them on new data.

                                                                                    People are working on this field, recent results suggest that continual learning can be possible by converting the input data to "LLMese"

                                                                                    • mochizou 11 hours ago

                                                                                      Maybe the practical path is to separate fast-changing memory from slow-changing weights. Most things an agent learns during use probably don't need to become parameters immediately.

                                                                                  • SerdarGl 20 hours ago

                                                                                    At massive scales 0th order methods will parallelize better than backprop especially along depth, u can train very deep models pipeline parallel without bubbles

                                                                                  • wg0 18 hours ago

                                                                                    It's computationally expensive and infeasible but what's the upside?

                                                                                    Genuine question due to unfamiliarity with the subject.

                                                                                    • alpu 14 hours ago

                                                                                      I think it can be useful for optical hardware

                                                                                    • ammarmalik17 13 hours ago

                                                                                      Bigger Dust models needing fewer population samples?

                                                                                      • oofbey 18 hours ago

                                                                                        This is super yawn-worthy. Instead of backprop for the exact gradient you can run forward passes a thousand times with perturbed weights and get a Monte Carlo estimate of the gradient. Not very clever. Extremely NOT useful.

                                                                                        But I guess the industry is littered with techniques for computing the same thing but vastly slower that some people find interesting. Homomorphic encryption. Zero knowledge proofs. Blockchain computing. Except in those cases there might be a legitimate reason to use it occasionally.

                                                                                        • ContinuityLab 14 hours ago

                                                                                          Moving beyond traditional backpropagation for transformer pretraining opens up fascinating avenues for alternative learning dynamics and architectural efficiency.

                                                                                          • vivzkestrel 16 hours ago

                                                                                            - stupid question from a neural network rookie

                                                                                            - isnt the whole point of back propoagation so that you dont guess weights by brute forcing them since that is computationally infeasible once you go beyond a dozen weights?

                                                                                            - if you dont use backpropogation, how exactly are the initial weights assigned them if they are not random values?

                                                                                            • AIorNot 18 hours ago

                                                                                              Hmm I was more interested in this alternative to backprop:

                                                                                              https://news.ycombinator.com/item?id=49701182

                                                                                              • wrecked_em 21 hours ago

                                                                                                Definitely more than meets the eye.

                                                                                                • whatever1 15 hours ago

                                                                                                  Remember kids, whoever is pitching space searching via complete enumeration just wants your wallet.

                                                                                                  • eriwang915 21 hours ago

                                                                                                    Dust's 243M model beating a 120x smaller one at most population sizes is the surprising part; bigger nets got more population-efficient, not less.