← Back to context

Comment by scosman

14 hours ago

Community effort happening here to build the ideal dataset: https://github.com/scosman/pelicans_riding_bicycles

Wow nice. If an llm could replicate these excellent examples, I'd consider the pelican benchmark fully saturated.

Come on, don't provide the smoking gun that shows how to draw a pelican riding a bicycle. If it's on the public internet it will end up in training data and invalidate this important LLM capability benchmark.