← Back to context

Comment by dozerly

2 months ago

I’m afraid you’re going to have to start randomizing your benchmarks somehow. I’m sure these models are trained on this problem by now.

If they did, they launched early! If they didn’t, their training data contains a bunch of poorly executed pelicans on bikes by other models.

Wait. Based on the results of the test linked above you think this model might have been trained to produce it? Did you look at the results?!

  • That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark.