← Back to context

Comment by Ariarule

1 day ago

Always glad to see more open-weight models, but this caption on the 2nd demo image had me do a double-take: "Land or Water Generalization Experiment: We recreated the viral X puzzle by asking Beam to create a fixed 180×90 grid for longitudes -179° to 179° and latitudes -89° to 89°, with 16,200 points. This puzzle is a few days old, so could not appear in the training data, thus testing the model’s generalization. Beam gets 95.5% coverage right, putting us between Opus 5 (92.5%) and Fable 5 (97.8%), which shows how well it generalizes to novel new tasks."

Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...

Yeah I remember when the original post about this came out. Def not recent. Though I think their point survives in that they didn't exactly RL on this.

I don't think age of the puzzle even matters, all models have search capacities these days

  • > all models have search capacities these days

    one would hope that they disable websearch and internet access (maybe all tools?) when doing generalization testing?

  • Model weights (what is being tested here) don't inherently "access the web" when inference is running. If the model has access to a web search tool, that's a different story.

  • Is search part of the model or the harness?

    • Maybe that's a rhetorical question but just in case - the search would always be part of the harness. A model is only handling next-token prediction for a given input. That token may be something like [[web search]] to invoke a tool call but the actual call would be handled by the harness.

      2 replies →

  • If they made that statement and knowingly had search enabled, it would essentially be fraudulent.