← Back to context

Comment by oliculipolicula

4 hours ago

Tweeter checked it with Astra. It seems like OAI could have pointed their own instance at it before launch. Because the source of the tip is likely someone at OAI, my guess is that they actually did check. But after the launch.

LLMs are unpredictably complex with potentially sigbificnaly different reaults based on random seed and seemingly insignificant prompt details) in the ideal case and nondeterministic in practice, so someone finding an error with a given LLM is not strong evidence that the result was not checked with an LLM, even with the very same LLM, previously.

  • No, not modern foundation models. This isn’t gpt-3.5-turbo. Although they are causal autoregressive, they have self consistency. You just have to verify multiple times to ensure you have averaged out any sampling errors.

I can throw Opus 5.5 at my code three times for code review and get three different sets of things it considers to be issues. I imagine all of them were checked with Astra at least once, but were they checked enough times?