Comment by simonw 2 months ago Here's a pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... 10 comments simonw Reply nikcub 2 months ago Real test here would be using tinker to fine tune a tinker model to generate pelicans calny 2 months ago well it's good they didn't train on the test! m3kw9 2 months ago How is this still a valid test if there is a good chance they will now specifically train on it? el_io 2 months ago Gist is returning 403 dozerly 2 months ago I’m afraid you’re going to have to start randomizing your benchmarks somehow. I’m sure these models are trained on this problem by now. tyre 2 months ago If they did, they launched early! If they didn’t, their training data contains a bunch of poorly executed pelicans on bikes by other models. tstrimple 2 months ago Wait. Based on the results of the test linked above you think this model might have been trained to produce it? Did you look at the results?! argee 2 months ago That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark. 0-_-0 2 months ago It's not a benchmark, it's a meme OrangeMusic 2 months ago To be fair, the "they're trained on this benchmark" response is also a meme.
nikcub 2 months ago Real test here would be using tinker to fine tune a tinker model to generate pelicans
m3kw9 2 months ago How is this still a valid test if there is a good chance they will now specifically train on it?
dozerly 2 months ago I’m afraid you’re going to have to start randomizing your benchmarks somehow. I’m sure these models are trained on this problem by now. tyre 2 months ago If they did, they launched early! If they didn’t, their training data contains a bunch of poorly executed pelicans on bikes by other models. tstrimple 2 months ago Wait. Based on the results of the test linked above you think this model might have been trained to produce it? Did you look at the results?! argee 2 months ago That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark. 0-_-0 2 months ago It's not a benchmark, it's a meme OrangeMusic 2 months ago To be fair, the "they're trained on this benchmark" response is also a meme.
tyre 2 months ago If they did, they launched early! If they didn’t, their training data contains a bunch of poorly executed pelicans on bikes by other models.
tstrimple 2 months ago Wait. Based on the results of the test linked above you think this model might have been trained to produce it? Did you look at the results?! argee 2 months ago That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark.
argee 2 months ago That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark.
0-_-0 2 months ago It's not a benchmark, it's a meme OrangeMusic 2 months ago To be fair, the "they're trained on this benchmark" response is also a meme.
OrangeMusic 2 months ago To be fair, the "they're trained on this benchmark" response is also a meme.
Real test here would be using tinker to fine tune a tinker model to generate pelicans
well it's good they didn't train on the test!
How is this still a valid test if there is a good chance they will now specifically train on it?
Gist is returning 403
I’m afraid you’re going to have to start randomizing your benchmarks somehow. I’m sure these models are trained on this problem by now.
If they did, they launched early! If they didn’t, their training data contains a bunch of poorly executed pelicans on bikes by other models.
Wait. Based on the results of the test linked above you think this model might have been trained to produce it? Did you look at the results?!
That’s the problem. If all (or zero) models were trained on it, it would be fine as a benchmark.
It's not a benchmark, it's a meme
To be fair, the "they're trained on this benchmark" response is also a meme.