← Back to context

Comment by jaggederest

4 days ago

The thing I like to do is to use models from different training sets - so for frontier, OpenAI criticizes Anthropic and vice versa. They're very much peanut butter and chocolate in that regard - I honestly can't be bothered to set up the whole MMLQUALA benchmark suites or anything, but I wonder how high "the two best models running at max thinking working together" would score compared to either individually.