← Back to context

Comment by cbg0

11 hours ago

But is the score really reflective of the quality or are both models benchmaxxing?

Both versions of DeepSWE (1.0 and 1.1) are likely not that meaningful anymore. Whether through models progression or through contamination.

Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD.

This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.