← Back to context

Comment by jascha_eng

9 hours ago

-5 on omniscience? https://artificialanalysis.ai/evaluations/omniscience

That's not particularly great.

That said I love that they don't seem to restrict cyber capabilities to any degree and even lean into it.

If the best model for cyber attacks is open for everyone to use it just makes us all safer I think. Of course you then also HAVE to use it or otherwise you're vulnerable, which is a great distribution play.

I ran it on my trivia game Redactle which features a redacted Wiki article. Mistral Large 4 is not very good. It can sometimes solve a game with ~40 guesses whereas the top models like Gemini 3.8 Flash or Grok 4.7 can one shot most puzzles. My benchmark here aligns with AA Omniscience. I also have a version where the text is rewritten to detect over fitting to exact wiki text which changes the scores but not the leaderboard order.

https://redactle.net/llm-leaderboard?view=vital-500

  • Gemini models are summarizing wikipedia articles all day (when being used for Google's ai answer), can we draw some conclusion from this, did they train it extra well on wikipedia content?

https://artificialanalysis.ai/models/mistral-large-4 for the main stats

                          Inte         Cost
                          llig          per
  Open Weight model       ence  Speed  Task

  Mimo-V.26-Pro            46     47   $0.13
  GLM-5.3 (max)            45     73   $2.01
  DeepSeek 4.1 Flash Max   39    227   $0.27
  Mistral Large 4 Preview  38    116   $1.13

  • Imo omniscience correlates better to how useful the model is in practice than the intelligence index. But you have to use both together of course.

    • Too late to edit, now including AA-omni Score [0] and 'by Domain' Software Engineering [1]. Also added comparison to the eyeballs-median of the 'top 10' models and then you see that indeed Mistral Large 4 Preview scores miserable in the AA-Omni indices.

                                Inte         Cost     AA-   Omni
                                llig          per    Omni  Softw
        Open Weight model       ence  Speed  Task   score    Eng
      
        Mimo-V.26-Pro            46     47   $0.13      8     33
        GLM-5.3 (max)            45     73   $2.01     14     37
        DeepSeek 4.1 Flash Max   39    227   $0.27     -5     34
        Mistral Large 4 Preview  38    116   $1.13     -5      5
      
        Closed/proprietary      ~50   110-  $1.50-    ~43    ~85
           median top 10               242   $7.50
      

      [0] https://artificialanalysis.ai/evaluations/omniscience [1] https://artificialanalysis.ai/evaluations/omniscience?detail...

>If the best model for cyber attacks is open for everyone to use it just makes us all safer

Issue is..

I don't believe for an instant that any of us, including US citizens, get access to the best models for cyber that the US has. I think any adversary would have to assume the models in use by the US side are unreleased.

US is not the only one dealing under the table by the way, I also think everyone should take China having unreleased models as an operating assumption at this point.

So Mistral is the best that the public gets access to. And that's if it's even the best? Benchmarks and pragmatic work have often been shown to be two radically different things in this industry.

If the best model for cyber attacks is open for everyone to use it just makes us all safer I think.

The NRA approach to AI safety.