Model weights (what is being tested here) don't inherently "access the web" when inference is running. If the model has access to a web search tool, that's a different story.
Maybe that's a rhetorical question but just in case - the search would always be part of the harness. A model is only handling next-token prediction for a given input. That token may be something like [[web search]] to invoke a tool call but the actual call would be handled by the harness.
Yeah it was rhetorical. Search would be implemented as a tool call. Pure intelligence tests would likely have limited tools. But maybe they would have a python sandbox to solve issues like Rs in strawberry.
> all models have search capacities these days
one would hope that they disable websearch and internet access (maybe all tools?) when doing generalization testing?
Model weights (what is being tested here) don't inherently "access the web" when inference is running. If the model has access to a web search tool, that's a different story.
Is search part of the model or the harness?
Maybe that's a rhetorical question but just in case - the search would always be part of the harness. A model is only handling next-token prediction for a given input. That token may be something like [[web search]] to invoke a tool call but the actual call would be handled by the harness.
Yeah it was rhetorical. Search would be implemented as a tool call. Pure intelligence tests would likely have limited tools. But maybe they would have a python sandbox to solve issues like Rs in strawberry.
If they made that statement and knowingly had search enabled, it would essentially be fraudulent.