← Back to context

Comment by cubefox

13 hours ago

> Derivative-free optimization can be useful for genuinely discontinuous objectives [1], but common neural network objectives are smooth and/or Lipschitz.

There is even an analog to the continuous derivative for discrete binary functions, called "Boolean variation": https://proceedings.neurips.cc/paper_files/paper/2024/hash/7...

Like for derivatives, there is a chain rule for Boolean variations, so you can use something like backpropagation, but without needing any expensive floating point math. Though I don't think this has been used much so far. There must be some other downside.