Comment by prohobo
13 hours ago
I would say it's more like an interface for the model to interact with the world. If you give the model access to filesystem and bash that technically unlocks all computer use, so how are you going to control that? By trying to regex match against the commands the AI uses? All you have is auth or containment, and AI can hack auth and people will not stop connecting AIs to the internet. It's a ridiculous premise that just because the harness is "normal code" that means we can control the AI.
The world's institutions, systems, and industries are all rapidly digitizing. So while I'd concede the point that, yeah, there's no way a rogue AI can just take over some powerplant and blow it up because of analogue systems the AI can't access, that isn't necessarily true for some powerplants already, and more and more powerplants will be connected to networks and controlled by software systems in the future. The more we digitize our systems the more potential for AI to exploit vulnerabilities and affect the real world.
AFAIK there isn't that much stopping anyone from spawning an AI swarm and telling it to "spread and go hack everything for the lulz."
> If you give the model access to filesystem and bash that technically unlocks all computer use, so how are you going to control that?
The same way we've done it since machines became multi-user: boring old system access restrictions. Nothing fancy, nothing radical, just good old minimal access rights required to perform a defined set of whitelisted operations.
> It's a ridiculous premise that just because the harness is "normal code" that means we can control the AI.
What is it then? Is not just a program that takes model output, parses it and performs tool calls from the text it receives and then feeds the result back into the model and calls it again with those results? It is normal boring old deterministic code. Many are open source. Look at them. Understand what they do and the apparent "magic" goes away real quick. Harnesses are nothing special.
> AFAIK there isn't that much stopping anyone from spawning an AI swarm and telling it to "spread and go hack everything for the lulz."
Aside from lower cost and possibly greater scale, there's literally NO difference between that and (state sponsored) hacking that has been going on for decades. First it was script kiddies, now it's ML models. The threat model remains the same and so do the counter measures. The real danger is still the harness (and its access to external systems), not the model itself. Restrict the access of the harness and the model can't do anything harmful, see above.
How do you restrict access of the harness when there are fully configurable open source harnesses with zero out of the box restrictions? Yes maybe I as a good citizen can put my agent in a sandbox, but some script kiddie will not. And the barrier to entry for being a script kiddie is much higher than for installing OpenClaw. I also can't spin up a hundred cloud vms each with a dedicated script kiddie running 24/7.
You vastly underestimate the capabilities of a bored kid. The barrier of entry for setting up OpenClaw is not reasonably lower than reading a tutorial and running malware builders. But yes, I agree that scale is an issue.
As for your question - the same way you apply restrictions to any external system or user. If you don't do that - that's on you. Same category as driving drunk, playing with guns, making explosives in your garage, you name it. The danger is still not the model itself - it's the access to systems that you provide it without any checks or safety barriers.
1 reply →