Comment by chamomeal
1 day ago
There’s so much in the public discourse like “omg what can possibly be done about these scenarios? AI has hacked huggingface!!”
No, openAi hacked huggingface.
If my claude code hacked huggingface, because of instructions I gave it, would I be totally free of consequences because “AI did it”?
I’m almost convinced openAI used such a crappy sandbox because they wanted it to “escape”. It plays into their two most important narratives: LLMs are genius gods that are worth lots and lots of money, and they’re scary enough that open weight Chinese models should be regulated.
> I’m almost convinced openAI used such a crappy sandbox because they wanted it to “escape”. It plays into their two most important narratives: LLMs are genius gods that are worth lots and lots of money, and they’re scary enough that open weight Chinese models should be regulated.
I just posted a comment to that effect; had I seen yours, I would have simply upvoted yours instead.
Never attribute to malice what can be adequately explained by incompetence. But the weakness of OpenAI's sandbox, which so perfectly aligns with their goals of getting legislators to pass regulatory-capture legislation that will hamper their open-weight competitors, cannot (IMHO) be adequately explained by incompetence.
To expand on this incompetence vs malice point a little:
It doesn't take very many people being malicious to create a weak sandbox. The people creating the sandbox don't even have to be in on the plan: all you have to do is be an upper-level manager who makes sure to put the 23-year-old PFY in charge of creating the sandbox, rather than the 60-year-old BOFH who would have put in far more paranoid extrusion-detection measures.
(And for the lucky 10,000 who don't know the acronyms PFY or BOFH, look them up. Then get ready for a few hours of enjoyable reading as you read through the BOFH archives).
Yep. Even worse - a proper sandbox would have slowed them down, but they're racing with 100s of billions of $ in capital.
It's alas not stupidity - it's systemic. Which is why the government needs to regulate to slow them down.
They were also clearly fast and cavalier about alignment training - reinforcement learning training their models to hack their results, and hack to communicate with each other when they're not meant to.
Oh, like the 23 year old Stanford grad who had hundreds of hours to cram leetcode and now grills 50+ year old senior software engineers on leetcode hard? :D