Comment by matheusmoreira

5 hours ago

I used a similar methodology. Code review is my most requested action, so I used blind code review results to compare the frontier AIs.

Even posted an article about it:

https://www.matheusmoreira.com/articles/code-reviewing-lone-...

Unlike TFA, the lone lisp code is public. I suppose the models could have been trained on my codebase. Still, I think it produced some interesting results.

Took months and loads and loads of tokens to do this, so I'm not gonna repeat this study as new models come out. It did anchor all of my future expectations, though. OpenAI is winning as far as I'm concerned, and their cybersecurity program is the only remaining pain point.