How specifically did misconfiguration lead to the HuggingFace attack? You could argue that its sandbox was misconfigured, sure. But suppose you had a similar incident where its intended task required access to the internet, and it veered off course in a similar manner. I don't think "misconfiguration" would be an accurate description of what went wrong in that hypothetical.
The doomers already have a term which fits pretty well: "AI misalignment".
there is a rogue agent - openai and the whole management chain from researcher to sama.
theres no separate agent, which is the point. the program might look like it, but that is an illusion of the interface. the llm produces text, and the harness executes commands based on text, based on what the human researcher included as things that can be executed
It does, indeed. Because OAI is not just a singular person, as in your scenario. No single person has access to controlling agents at the scale OAI has. Let's not conflate Frontier providers with "Person".
I'm not exactly sure why you think this distinction is so important. I think my point stands if you replace "Person" with "OpenAI". In any case, I presume the swarms OpenAI has been researching will be available to the general public before too long.
What label do you prefer for the agent that did something it was not asked to do?
bot. and we even have a word for program not behaving the way the way it was intended.
"Bot" doesn't carry any implication of unintended behavior. You could call it a "buggy" bot, but these aren't ordinary software bugs.
There's no simple bugfix which will address AI misalignment. It's essentially been an open research problem for upwards of a decade.
10 replies →
Unreliable computer program.
Most unreliable computer programs won't launch research programs consisting of thousands of pages of text to find creative ways around obstacles.
Misconfiguration. We deal with lots of applications every day that can do terrible things if you get the config slightly wrong.
say.. https://www.investor.gov/introduction-investing/investing-ba...
How specifically did misconfiguration lead to the HuggingFace attack? You could argue that its sandbox was misconfigured, sure. But suppose you had a similar incident where its intended task required access to the internet, and it veered off course in a similar manner. I don't think "misconfiguration" would be an accurate description of what went wrong in that hypothetical.
The doomers already have a term which fits pretty well: "AI misalignment".
2 replies →
there is a rogue agent - openai and the whole management chain from researcher to sama.
theres no separate agent, which is the point. the program might look like it, but that is an illusion of the interface. the llm produces text, and the harness executes commands based on text, based on what the human researcher included as things that can be executed
Person: "AI, please make me paperclips."
AI: "OK, I've now converted the entire planet into paperclips."
Alien observer #1: "Wow, that was a rogue AI!"
Alien observer #2: "False. We need to place the blame where it belongs, on the person who requested the paperclips."
Ultimately this type of terminology dispute has a tendency to miss the point.
It does, indeed. Because OAI is not just a singular person, as in your scenario. No single person has access to controlling agents at the scale OAI has. Let's not conflate Frontier providers with "Person".
I'm not exactly sure why you think this distinction is so important. I think my point stands if you replace "Person" with "OpenAI". In any case, I presume the swarms OpenAI has been researching will be available to the general public before too long.
3 replies →
That legal fiction works both ways.
3 replies →