Comment by pampas
1 hour ago
I ran it on my trivia game Redactle which features a redacted Wiki article. Mistral Large 4 is not very good. It can sometimes solve a game with ~40 guesses whereas the top models like Gemini 3.8 Flash or Grok 4.7 can one shot most puzzles. My benchmark here aligns with AA Omniscience. I also have a version where the text is rewritten to detect over fitting to exact wiki text which changes the scores but not the leaderboard order.
Gemini models are summarizing wikipedia articles all day (when being used for Google's ai answer), can we draw some conclusion from this, did they train it extra well on wikipedia content?