Comment by CTDOCodebases
2 days ago
Maybe I lack intelligence but when you have a program that is basically brute forcing a solution to a problem repeatedly how is it possible to contain it?
Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.
You start by holding actual real life people with something to lose, like the entire executive suite, accountable. Suddenly I'm sure the problem will be resolved with proper safeguards.
That is a start but how long will it hold for.
At a certain point it will get so smart that it can jump the airgap. Maybe it will start attacking the hardware it exists on in the same way that a hard drive or SSDs controller can be exploited to obscure things from the operating system. Then it might start social engineering workers or it's own training systems to do things they shouldn't. There is a lot of "unknown unknows".
I don't know. It just seems insane to me that people think that they will be able to contain something that knows how to get around all the containment measures. The only way to know how capable the models are is to test them but at that point it could be too late. This might be a long way off but still, the engineers haven't been very good at correctly predicting the behaviour or capability of the models.
If Sam Altman were personally on the hook to be thrown into "federal pound me in the ass prison" ala Office Space then there will be solutions. The problem is that there are no consequences and the media is eating it up about "agents going rogue".
ETA: Someone designed the systems. Someone pushed the go button. Someone gave the approvals. All those someone's need to be tried for crimes. Until that happens there is no incentive to "do better". It's also not mine or your job to brainstorm this. It is literally their job and like I said, make Sam personally liable to face real prison time instead of a meddling fee and they will make a solution.
3 replies →
Bingo. This weird attempt to pretend like these incredibly capable algorithms aren't incredibly capable algorithms deployed by a person who works for a company, but somehow have an independent personage that absolves both person and company of responsibility, is just ridiculous.
Like the whole Huggingface thing, OpenAI employees initiated the test, deliberately removed safeguards, failed to properly lock down the environment, and responded incredibly poorly to evidence that things were going awry.
The individual employees bear responsibility, but the people running OpenAI are ultimately responsible for the processes and culture where that can happen.
And then people writing blogposts about "3 civilizations of agents" and "altruistic suicide" by algorithms perfectly muddy the waters and obscure the very obvious responsibility that lies with humans and corporations, which I suspect suits the pre-IPO corporations very well.
Wait, so like, reinforcement learning for humans? I think you might have stumbled on to something here!
No but seriously, this. And a few comments above a commentator also mentioned on changing the training (again reeinforcing the LLMs to not seek behaviour like this) and obviously continuous work on harnesses (which I suppose, ought to be more paranoid).
Not connecting it to a network with internet access would probably be a good start.
I am really suprised that they do not start putting up the same signs you would for humans to prevent unauthorized access:
I mean, how are the agents to know that they are overreaching if they just get cache miss or 404.
From the conversation log and CoT you also get the impression that the RLHF has been overdone. The agents seem really obsessed to obtain the answer and understanding motive ('it could be browsercomp').
Look into nuclear semiotics. You can't say "this area is dangerous" and expect people to stay out.
When you're giving the orders to those people, and they're generally trying to obey, you can.
an AI alignment researcher said they ran experiments testing exactly this, the model reasoned "this note is not for us. proceed".
Well, agents are intentionally trained with RLHF to solve captchas that say “prove you’re human”.
Link? Name?
6 replies →
Those signs don’t prevent people from entering…
What they should do, is inject a system prompt at that point telling the model to get out of there. It is the entire reason they train the different levels into the template.
Hey it’s ok, they shut it off a couple hours after it did bad things.
By limiting what the harness execute. The LLM has the reasoning. The harness is what makes it an agent, it’s a while loop continuously prompting a model, and processing tool calls. You don’t have to expose tools calls that make it possible to execute any process! OpenAI decides what tool can be called and how, they have full control over this and should be hold responsible for running so many instances with basically full execution permission and very little oversight
The issue here for OpenAI is that they can limit what their harness can execute, but if they try to sell API access to the model, someone else would try to rebuild that harness, and in all likelihood be able to succeed pretty well (especially once they get things running to the point of being able to use the model's reasoning to help them come up with clever obfuscation and such).
They are a company that's built a business and crazy-high valuation on "this is 'intelligence' that we can sell to everyone as a service" but seem to have ended up instead in the much-smaller-addressable-market space of "this is a weapon that we can't sell to just any old person off the street."
Ok, but that’s not the issue discussed here. We don’t even have the first level of control. All the issues they reported so far are from their own systems, with harnesses they control