← Back to context

Comment by syntacticsalt

18 hours ago

I'm skeptical as to whether zeroth-order methods really lend themselves to a Bitter Lesson argument. First-order methods don't explore the loss landscape optimally, but the loss function tends to be nonconvex, and zeroth-order methods don't address that issue head on. Dust smooths, and so do applicable first-order methods. Remove the nonconvexity issue, and I suspect Dust's purported advantages evaporate (based on published theoretical work), so it's pretty odd to me that the paper never discusses convexity.

I could buy that this method scales better than previous zeroth-order methods, and that's interesting, but it doesn't seem like enough of a moat to keep improved first-order methods from drinking its milkshake, except in cases where a zeroth-order method is already a primary option: the network needs to call a simulator that doesn't expose gradient-like information. (In cases where gradients don't exist, I'd still argue for other options, e.g., Clarke-generalized gradients where applicable, so long as those can be computed with the available information. I know this technology has been published for automatic differentiation, so I would imagine it could be incorporated into backprop and used with a suitable optimization algorithm.)

Is it fair to define Reinforcement Learning (RL) as forcing a gradient onto a system / simulator?

Though, I suppose RL has non-gradient based methods too.

Orders-of-magnitude improvements in compute efficiency are needed to become a practical replacement for backprop… but those improvements are coming.

As a mixture, could activation-space search produce useful teaching targets for backprop?

Zeroth-order search would discover candidates, first-order learning would consolidate them. The potentially valuable step is converting a sparse judgment into a reusable training target.

This also changes the relevance of convexity.

  • I'm not sure I follow why your argument changes the relevance of convexity? If the problems were convex, I don't think we'd be having this discussion -- first-order methods would tend to win.