← Back to context

Comment by simonw

2 months ago

Here's a pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

Real test here would be using tinker to fine tune a tinker model to generate pelicans

How is this still a valid test if there is a good chance they will now specifically train on it?

I’m afraid you’re going to have to start randomizing your benchmarks somehow. I’m sure these models are trained on this problem by now.

  • If they did, they launched early! If they didn’t, their training data contains a bunch of poorly executed pelicans on bikes by other models.

  • Wait. Based on the results of the test linked above you think this model might have been trained to produce it? Did you look at the results?!

    • That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark.