← Back to context

Comment by emp17344

10 hours ago

Or it’s an indication that progress has plateaued. But instead of accepting this, you’d rather we just throw out the entire benchmark.

Why would you accept it when the benchmark's ranking is obviously nonsense. It literally has muse spark 1.3 above 6 astra, 5.6 sol and fable 5. Anyone who has played with any of these models for any amount of time would immediately realize that this is total bunk.