Comment by vmg12
3 hours ago
Models have been good at finding exploits for half a year now, this is not how we found out LLMs were good at hacking, you are rewriting history.
We knew models much weaker than mythos were good at hacking the problem they had was that when finding exploits they had too many false positives.
Either way, putting artifactory on the sandbox security boundary is obscene negligence. There is no reason to believe artifactory is secure.
If you listen to the OpenAI Black Hat talk it is very obvious they were surprised at the level of capability on display and felt it was novel.
But I guess OpenAI's security researchers acting surprised is part of some grand conspiracy to manage PR?
It may have been novel for OpenAI but we already had Mythos at this point and this talk by Nicholas Carlini.
https://www.youtube.com/watch?v=1sd26pWhfmg
We already knew LLMs were capable of finding exploits like this.
The Mythos issue is different though. It did not use any novel exploits. Irregular, the vendor they were using, misconfigured the environment to allow internet. Every other AI company also uses Irregular and that's why we saw so many articles come out at once.
The OpenAI HF incident is separate from this. It involved actual zero days, teamwork and message passing, and sophisticated chains of exploits.
2 replies →