← Back to context

Comment by usernametaken29

20 hours ago

> There are many interesting open questions. The first is whether, and how, Dust can find better directions than backprop’s first-order gradient

Both algorithms are bound by the same Pareto frontier based on the Empirical Risk Minimisation Principle, so they’re already on the same trajectory. Interestingly backprop is limited by conditioning of the Hessian matrix in order to converge (differentiate correctly). So removing this limitation is actually a great step. I’m excited to see a comeback of evolutionary methods because they’re much more general, albeit costly and naive. We’re now very close to what can be described best as brute forcing the Pareto frontier out of our datasets. Not sure that’s what we want but I have no better ideas either.

> Interestingly backprop is limited by conditioning of the Hessian matrix in order to converge (differentiate correctly)

What does this mean?