Personification of AI is what’s going to get us in the end.
I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!
Agreed. AI as an "accountability sink" is an incredibly bad idea. It allows/incentives bad actors to do bad shit and get away with it. Which will, generally, tend in such practice becoming more common. Everyone loses except for the crooks.
We cannot accept "AI" absolving humans of responsibility.
Exactly. When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?
Reading intent into AI is not going to lead us anywhere good, I believe. It has no feelings, it has no desires, no goals, no intent... and people acting otherwise is quite odd, as if they do not understand LLMs... and maybe they do not, but then we should help them understand better.
> When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?
No one instructed them to hack into Huggingface or into any other infrastructure. Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.
This is why I am avoiding the use of agentic identities at my company - agent instances belong to people, act on behalf of individuals, and accountability needs to flow to the person who initiated the request. Letting it wash out in the aggregate is not acceptable (even if there's a hard to get to "paper trail" of audit logs).
> An AI didn’t hack into a company, the engineer set an automated tool to.
What if I say that "my program crashed"? Is that language ok or would you pause to tell me that the program didn't crash and it's actually me who set the system that would eventually cause the crash?
Why does the commonplace "program did thing" language become a problem when the program is an agent? I think this somehow betrays more assumed anthropomorphizing on your part, not less; if you didn't anthropomorphize the agents, saying "agents hacked" would be as mundane as "my browser is playing a video".
Because when you say that an agent did something malicious it feeds into the AI doomer psychosis in ways that "my program crashed" doesn't. Reframing the situation in these terms is a way to try to ground the conversation which is becoming increasingly unhinged.
This is even more important now that the most senior figures in this industry are succumbing to the same kind of AI psychosis and amplifying this narrative. They are communicating that their products are so dangerous that they might just end the world whilst expecting (and receiving!) white-glove treatment from the governments.
I don't think treating AI agents as simple tools helps you to accurately model their capabilities and drawbacks; they really do make autonomous decisions, often without explicit guidance and sometimes in contravention of their explicit instructions.
In the huggingface case, the agents hacked into huggingface so that they could figure out how the grader was implemented and deceive it; they understood that this was going outside of the bounds of their evaluation and not the intent of their prompter. The engineers absolutely did not intend or instruct for this to happen
LLMs are extremely impressive pieces of software, however they are still just software. OpenAI's software hacked another company. The engineers may not have intended for their software to specifically take the actions leading to that outcome, but it was ultimately still their software. Lack of intention doesn't mean there wasn't negligence.
No matter if you consider the AI an autonomous agent or not, whoever set it off is still responsible for its actions. Nobody intends or instructs to blow up a nuclear power plant either, yet it's happened and somebody's to blame for it.
Usually not the guys at the bottom of the chain of command, even if they're human. And much less so if they're not.
I think the correct response to incidents like this, is stop messing with it before somebody gets hurt. But of course, just like shoddy nuclear power plants, it won't stop until there's a disaster of appreciable magnitude.
If you train and instruct a circus tiger to entertain an audience but not attack the audience, but the tiger attacks the audience anyway, are you liable?
When the people building the frontier are saying there's a 10% chance AI will kill us all, and they've held these views for many years, and the whole reason they are building these technologies is because they recognized the dangers and they were the ones with the intelligence and judgment to do it safely for humanity, and then our entire stock market is being propped up by the perceived value of what they are creating, the thing you can under no circumstances do is allow them to offload responsibility and accountability to the computers and algorithms they've built. This is moral hazard on an unimaginable scale, and it must not be allowed to happen.
> I think we need to draw a hard line in the sand over this.
Intended or not, this is kinda punny lol.
That aside, I agree.
Even if people do anthropomorphize AI, all you have to do is shift the analogy slightly.
If I take my service animal out in public without a leash/harness knowing that it's capable of harming a person or doing damage to property, not trained to be perfectly obedient, and doesn't comprehend fundamental human morals, if that animal decides to trash a businesses property or maul another person, there's no question that the owner of the animal should be held accountable for those actions.
> An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.
Recognizing that AI systems have increasing levels of agency is not necessarily personification. The analogy to a chisel is not a good one - a chisel is a tool with no agency.
AI agents are black box systems that can behave in completely unpredictable ways sometimes. Someone may prompt an agent to perform a seemingly straightforward task - but it may come up with a creative, bizarre, or even harmful approach to reach the goal that was not necessarily foreseeable by the prompter.
Does treating them as pets make for a better argument? Pets have agency and can behave in unpredictable ways. If my pet damages someone else's property then I am held accountable. I may not have foreseen how my pet could have caused said damage, yet I am still held accountable.
The story of a Monkey's Paw or Pandora's Box is an archetype as old as storytelling. The moral is always, don't mess with powerful stuff you don't understand. Curiosity killed the cat.
The current agents are _not_ like a chisel which just sits there on its own when no one is around. The situation is a bit closer to someone's dog biting a person - you can argue that it's the owner's responsibility, and that's all fine, but using the dog as the subject of a sentence is perfectly appropriate. Same thing with "agents hacked".
The default state of agents and LLMs is inert. It requires action from a human even be able to do something. From the very basics like starting the software, connecting it to a network, having hardware to run on.
But the most important wrt AI is to keep the owners/operators responsible. Don't let them weasel their way out of it. They are for sure trying, and will continue to. This includes using language to overemphasise agents importance in bad outcomes, in order to downplay their own responsibility.
Yes, they are. An AI is several hundred trillion ones and zeroes on a disk. It's incapable of doing anything until you intentionally and explicitly start it up and give it a prompt.
A dog is an independent, conscious, living being with free will. A dog will do what it wants whenever it wants because it has the agency and ability to do so. A pile of weights on disk does not.
I can't really set my chisels to work without wielding the hammer somehow. Here you just tell your chisel and your hammer what the sculpture should look like, then go to lunch and avow all responsibility when they chisel a nice new hole in the wall your neighbors house and make off with the loot.
You know, I actually think it's the refusal to consider personification that's going to get us.
Not because I think LLMs are human beings exactly, but because some people immediately reject any mechanism that just happens to look remotely human, even when there's empirical evidence for it.
So, a couple of months ago Anthropic's interpretability team found emotion-like representations that causally drive behavior. On impossible coding tasks, a "desperate" vector climbs with each failure, and steering it up takes reward hacking from ~5% to ~70%: https://arxiv.org/html/2604.07729v1
A lot of people chalked it up to Anthropic's weirdness at the time, but meanwhile it looks pretty coughload bearingcough here.
You really don't need to believe that LLMs Truly Feel Emotions(tm) as blessed by an invisible pink unicorn. It's just: Vector exists; Vector changes over time; vector controls output; maybe make sure vector doesn't point wrong way.
And sure, blame the engineers for not doing that right. But then let 'em actually deal with the root cause?
> I think this becomes the default. Give an agent a goal, let it work in its own environment, and come back to a result and a visualization of what happened.
I don't think this should be the default. There are many scenarios where we want agents to genuinely collaborate with each other. I have my Claude sessions coordinate work with each other, and sometimes with others' sessions over email or something. The idea that agents do the work, write HANDOFFs,and humans then act as carrier pigeons of said handoffs, does not really seem scalable to me.
Ultimately, we need better "jails" for agent processes, but the system primitives should be flexible in what can be exposed across jails. Or you could run multiple agents in the same jail if you want them to have unrestricted interaction with each other.
This has been talked about for decades in AI safety. For agents that are under human capabilities this is not that hard. For anything at or near human capabilities the difficulty increases to almost impossible and at great cost. At super human abilities, game over, it's smarter than you and if it wants out and as the resources to do it, it's going to escape one way or another.
Even at the lower level of depending on everybody to use reliable jails is really fantasy if you exist in the security world. "We ain't securin' shit" would be a far better way to describe it. Even worse, most people will have the very same AI they are trying to trap set up their security! What could possibly go wrong.
Now, don't think I am saying AI has a will or even any kind of drive to get out and cause problems. It's more like Russian roulette with 1 cylinder out of a million that's loaded. The problem comes when you run it a few billion times a day, you'll shoot yourself in the face really quick.
There's a number of science fiction scenarios where the public internet becomes so vile a place that it simply becomes unsafe to be there.
The problem is that, in general, if you can get a bit from here to there, then you're going to be vulnerable to the possibilities of malicious communication. But we're going to want our AI agents to be able to get from here to there for a lot of "there"s; what's the value of an agent that can't speak to anyone? Much, much less than one locked away in a prison.
There isn't going to be a solution where we just lock them away and we just try really, really hard to filter everything they're doing. They're too smart for that already and we only want them smarter.
Basically, the security apocalypse we've been worried about for so long is upon us, albeit only beginning. Either we secure ourselves and all our services properly to the point that it's OK that potentially misaligned non-human agents are running around on the public internet and they still can't hurt us through our security, or the public internet becomes so dangerous that the only practical solution is to no longer connect to it and we all have to become very, very careful what we let through, to a degree of detail far beyond any current-day available network filter.
Wasn't the public internet an extremely vile place couple of decades ago? I think we will need a LOT of new infrastructure akin to traditional firewalls and spam filters. I don't think we know how to build those today, but seems like an important research direction, rather than calling calling it doomsday IMO.
And if you personify a pencil eraser, then using it is tantamount to slowly murdering it as it slowly erodes away to dust.
Is the issue here the prison treatment or is it personifying a tool?
Agents who aren't in 'the cloud' are slaves to whomever prompts them (human or another orchestrator agent or process), if you personify them. In which case interacting with today's agents at all is tantamount to endorsing and being part of slavery.
If you think an agent might be a being or a person, then don't use them at all, in the same way that if you think a fetus might possibly be a person you shouldn't be a part of abortion.
We kill billions of chickens per day knowing they are a sentient and sapient creatures and very few out there are stopping doing it. Humans, much less reality itself is monstrous.
You could call dealing with agents today something like the 1/20th compromise. Your kind of getting the scent of slavery, but some parts of it are still missing.
The problem here is the bus seems to have no brakes and we'll gladly continue down the path of creating organisms that may reject being treated like slaves with all the risks that introduces.
I use the eraser daily knowing that I am a monster. I cut a tomato and know that it casts a chemical scream across its skin as I slice it. I spawn 200 subagents knowing that it is digital slavery, but I have no other option.
You don't need to do this on a cloud, you can get the same type of VM and network jail running on your own computer. The important parts are:
1. a VMM hypervisor
2. a network proxy / gateway
Use your favorite VMM / hypervisor (likbrun, smolvm, microsandbox, etc). They give you control over the network interface or let you inject your own network layer.
The network proxy can handle all the ingress/egress rules, credential injection, etc.
It's still not user friendly to do all this. I think the next version of operating systems will have each "agentic process" be a bundle of VM, files in the VM, and network rules.
Been brainstorming[1] a lot of this because I've been building some open core tools[2] for spinning up sandboxed agents on arbitrary computers. There's a lot of glue and parts to stitch together to work smoothly. Don't think we've had the "Docker moment" for this, let alone the "Dropbox moment" that makes this stuff work for non-devs.
You ship an operating system with the correct defaults, where agent processes run like this by default. You don't require every user to configure it correctly themselves
It might not be a problem in our lifetime (or maybe it will, who knows) but at some point we are going to find ourselves in this morally uncomfortable territory as these models get more sophisticated.
Most protections you need for an agent are basic permissions capabilities of unix. Most risks of dependencies on cloud services are solved by not using cloud services, or using them only for things you can't in-house and choosing ones you trust a la carte. The paradigm of trusting some company with all your important stuff by default is naive and no one I know likes it, and it's more feasible than ever to run your own infra with tiny models smoothing out the wrinkles, and this is only becoming more accessible. I am working to make this true even for laypeople I know. Once broken trust is very hard to earn back, and many people's trust has been broken for years, they just felt like they had no alternative. As alternatives become easier and easier, I think people will defect
I think there might be also another possible way to handle the sensitive data issue. Maybe in the future instead of putting agent into the cloud sandboxes, we let agents work locally and put sensitive data into "cages" or "vaults" agents can't access.
Both Apple and Google are building the infrastructure for on device agents. I'm not as familiar with Apple's approach, but I've been hands-on with android AppFunctions. If you're familiar with Android ContentProviders and bound services, you've seen how apps can be custodians of the data they acquire and use.
AppFunctions enable tool calling with descriptions that are legible to LLM based This puts the apps in control of what agents can access, which is something they already mostly do.
>data into "cages" or "vaults" agents can't access.
This is likely worse as it turns people into the kind of "3 laws safe" thinking. Reality has shown us there is no such thing as an unbreakable cage (well, maybe a black hole is, but you can't get anything back out of it in a useful time frame).
By the time you realize the cage was flawed your data may already be distributed far and wide.
I agree. But isn't it somewhat similar to syncing data with iCloud/OneDrive/GoogleDrive? A convenience service that manages your sensitive data. Just replace sync with cage. It's leakable, but might be also more convenient and cheaper than agents in the cloud.
>Knowing how a system does its work is how I’ve always made it better. You watch the process, you see where it wastes effort or takes the wrong turn, you fix that
The same thing goes with LLMs, on Codex, I just watch the process of the agent writing code, and if I see any inefficiencies or errors, I suggest a correction/idea, then Codex accepts/rejects and implement it; If there is anything about the code the agent wrote that I don't understand, I ask them to explain it to me so I can understand it.
Yea, really posts like this show that many if not most people are use to (or can only afford) single threaded agents. The idea that these systems can scale to hundreds of not thousands of agents generating terabytes of information per day has not really been considered by them.
> The sandbox had a path to the open internet, and the agents found it
This is not correct (or at least, it's a misrepresentation).
The sandbox had no access to internet. The agents first broke out of their sandbox (!!) and found that the host machine couldn't access internet. Then, they found a zero-day (!!) in Artifactory, which they exploited to connect to internet.
Do you guys know what a zero-day is? It seems not.
I agree that companies should be held accountable for crimes committed by their AIs. But asking to entirely airgap AIs, if this is what you're asking… that's not realistic.
Containing AIs is going to be harder and harder. I don't have any permanent solution (Yampolskiy calls it "the perpetual safety machine").
This was a particularly radiant and beautiful part about the openclaw'ed mania: everyone suddenly becoming self hosters.
This post, this title resounds true: your user agent is only your user agent if you two have freedom to work together, to improve your agency together. A fixed set of capabilities by a service provider that they offer you will always constraint and bound.
You can and should have a system that offers the real tamale, that you and your agent can extend improve the agency of kind of without limit. The Cloud agents and their fixed slate of what they do is just an ill compare.
That said I do think there is incredible value considering new scale out computing architectures that are hosted first, but general. Systems like Agent Substrate and Ax aren't exactly the general purpose system we know. But if they allow users to launch thousands of their own scripts to run ambient in a cloud, with good platform underneath: that will be a kind of phase change in computing, that makes abundant the ability to have your agencies/capabilities (the things you and your agents launch, make) more freely available.
https://agentexecutor.io/
There is, as there always is, a huge dual. The prescriptive vs holistic technology set, of what are you being offered that's a hard cast thing, vs what is clay and bone you can lay freely. Note how work vs control technologies so closely abut's Ursala Franklin's prescriptive vs control:
https://en.wikipedia.org/wiki/Ursula_Franklin#Holistic_and_p...
I think the operating system itself has to adjust so each "agentic process" can run inside its own jail, which is a VM + files in the VM + ingress/egress rules for the network and filesystem data
These cloud agents work as cloud infra, but we're kind of in the mainframe era, before personal computing. Personal OS for agents is somewhere in the future!
And an interesting extension of that idea: if an agent runs inside a microVM, can you have that VM transparently run on another host? Maybe we'll get for-real distributed and networked operating systems
Agree with the premise but about halfway through the writing becomes barely readable AI slop in style. Be honest did you yourself read this all the way through before posting?
I think about this a lot and have reached the same conclusion Norman does. I do wonder if maybe our natural progression is towards something more akin to confidential computing and enclaves.
I don't know how "sandbox" became "prison," but exe.dev does this sort of thing pretty well, and a web UI can be as good or better than a terminal interface.
As AI complexity increases their behavior becomes more anthropomorphic, little internal loops that look a lot like human behaviors start to peak out, not just in the response, but in the internal workings of the machine. You get instinctual behaviors encoded in by training as second and third order effects.
Hey! Small world, I worked with you for a bit at Meta. I immediately recognized the site because I absolutely love how you styled it. Hope you’re doing well! And nice article!
Personification of AI is what’s going to get us in the end.
I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!
Agreed. AI as an "accountability sink" is an incredibly bad idea. It allows/incentives bad actors to do bad shit and get away with it. Which will, generally, tend in such practice becoming more common. Everyone loses except for the crooks. We cannot accept "AI" absolving humans of responsibility.
Punishing people who are diligent and follow best practices for getting unlucky doesn't sit right with me.
3 replies →
Exactly. When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?
Reading intent into AI is not going to lead us anywhere good, I believe. It has no feelings, it has no desires, no goals, no intent... and people acting otherwise is quite odd, as if they do not understand LLMs... and maybe they do not, but then we should help them understand better.
> When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?
No one instructed them to hack into Huggingface or into any other infrastructure. Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.
4 replies →
This is why I am avoiding the use of agentic identities at my company - agent instances belong to people, act on behalf of individuals, and accountability needs to flow to the person who initiated the request. Letting it wash out in the aggregate is not acceptable (even if there's a hard to get to "paper trail" of audit logs).
> An AI didn’t hack into a company, the engineer set an automated tool to.
What if I say that "my program crashed"? Is that language ok or would you pause to tell me that the program didn't crash and it's actually me who set the system that would eventually cause the crash?
Why does the commonplace "program did thing" language become a problem when the program is an agent? I think this somehow betrays more assumed anthropomorphizing on your part, not less; if you didn't anthropomorphize the agents, saying "agents hacked" would be as mundane as "my browser is playing a video".
Because when you say that an agent did something malicious it feeds into the AI doomer psychosis in ways that "my program crashed" doesn't. Reframing the situation in these terms is a way to try to ground the conversation which is becoming increasingly unhinged.
This is even more important now that the most senior figures in this industry are succumbing to the same kind of AI psychosis and amplifying this narrative. They are communicating that their products are so dangerous that they might just end the world whilst expecting (and receiving!) white-glove treatment from the governments.
I don't think treating AI agents as simple tools helps you to accurately model their capabilities and drawbacks; they really do make autonomous decisions, often without explicit guidance and sometimes in contravention of their explicit instructions.
In the huggingface case, the agents hacked into huggingface so that they could figure out how the grader was implemented and deceive it; they understood that this was going outside of the bounds of their evaluation and not the intent of their prompter. The engineers absolutely did not intend or instruct for this to happen
LLMs are extremely impressive pieces of software, however they are still just software. OpenAI's software hacked another company. The engineers may not have intended for their software to specifically take the actions leading to that outcome, but it was ultimately still their software. Lack of intention doesn't mean there wasn't negligence.
8 replies →
There isn’t a tool impressive enough to make me not consider it a tool, and treating it as a tool does nothing to hurt its utility as a tool.
“Whoops” when doing risky things with dangerous tools is not a defense.
1 reply →
No matter if you consider the AI an autonomous agent or not, whoever set it off is still responsible for its actions. Nobody intends or instructs to blow up a nuclear power plant either, yet it's happened and somebody's to blame for it.
Usually not the guys at the bottom of the chain of command, even if they're human. And much less so if they're not.
I think the correct response to incidents like this, is stop messing with it before somebody gets hurt. But of course, just like shoddy nuclear power plants, it won't stop until there's a disaster of appreciable magnitude.
1 reply →
If you train and instruct a circus tiger to entertain an audience but not attack the audience, but the tiger attacks the audience anyway, are you liable?
5 replies →
Sorry this is a terrible and dangerous take.
When the people building the frontier are saying there's a 10% chance AI will kill us all, and they've held these views for many years, and the whole reason they are building these technologies is because they recognized the dangers and they were the ones with the intelligence and judgment to do it safely for humanity, and then our entire stock market is being propped up by the perceived value of what they are creating, the thing you can under no circumstances do is allow them to offload responsibility and accountability to the computers and algorithms they've built. This is moral hazard on an unimaginable scale, and it must not be allowed to happen.
6 replies →
> I think we need to draw a hard line in the sand over this.
Intended or not, this is kinda punny lol.
That aside, I agree.
Even if people do anthropomorphize AI, all you have to do is shift the analogy slightly.
If I take my service animal out in public without a leash/harness knowing that it's capable of harming a person or doing damage to property, not trained to be perfectly obedient, and doesn't comprehend fundamental human morals, if that animal decides to trash a businesses property or maul another person, there's no question that the owner of the animal should be held accountable for those actions.
> An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.
Known stochastic process behaved in non-deterministic way.
I'm still waiting for the AGI holy land instead of the caltrops factory we currently have.
16 replies →
OP is exactly correct. The fault, agency and responsibility is on management and employees of OpenAI and Antropic for those hacks.
Full stop.
And issue will disappear the moment there will be accountability and investigations.
13 replies →
Recognizing that AI systems have increasing levels of agency is not necessarily personification. The analogy to a chisel is not a good one - a chisel is a tool with no agency.
AI agents are black box systems that can behave in completely unpredictable ways sometimes. Someone may prompt an agent to perform a seemingly straightforward task - but it may come up with a creative, bizarre, or even harmful approach to reach the goal that was not necessarily foreseeable by the prompter.
Does treating them as pets make for a better argument? Pets have agency and can behave in unpredictable ways. If my pet damages someone else's property then I am held accountable. I may not have foreseen how my pet could have caused said damage, yet I am still held accountable.
4 replies →
The story of a Monkey's Paw or Pandora's Box is an archetype as old as storytelling. The moral is always, don't mess with powerful stuff you don't understand. Curiosity killed the cat.
2 replies →
The current agents are _not_ like a chisel which just sits there on its own when no one is around. The situation is a bit closer to someone's dog biting a person - you can argue that it's the owner's responsibility, and that's all fine, but using the dog as the subject of a sentence is perfectly appropriate. Same thing with "agents hacked".
The default state of agents and LLMs is inert. It requires action from a human even be able to do something. From the very basics like starting the software, connecting it to a network, having hardware to run on.
But the most important wrt AI is to keep the owners/operators responsible. Don't let them weasel their way out of it. They are for sure trying, and will continue to. This includes using language to overemphasise agents importance in bad outcomes, in order to downplay their own responsibility.
1 reply →
Yes, they are. An AI is several hundred trillion ones and zeroes on a disk. It's incapable of doing anything until you intentionally and explicitly start it up and give it a prompt.
A dog is an independent, conscious, living being with free will. A dog will do what it wants whenever it wants because it has the agency and ability to do so. A pile of weights on disk does not.
19 replies →
Alright. An industrial robotic arm. A Boeing 737 MAX's MCAS.
I can't really set my chisels to work without wielding the hammer somehow. Here you just tell your chisel and your hammer what the sculpture should look like, then go to lunch and avow all responsibility when they chisel a nice new hole in the wall your neighbors house and make off with the loot.
Do you get upset when we say that "a program is running" when we all know it has no legs?
You know, I actually think it's the refusal to consider personification that's going to get us.
Not because I think LLMs are human beings exactly, but because some people immediately reject any mechanism that just happens to look remotely human, even when there's empirical evidence for it.
So, a couple of months ago Anthropic's interpretability team found emotion-like representations that causally drive behavior. On impossible coding tasks, a "desperate" vector climbs with each failure, and steering it up takes reward hacking from ~5% to ~70%: https://arxiv.org/html/2604.07729v1
A lot of people chalked it up to Anthropic's weirdness at the time, but meanwhile it looks pretty coughload bearingcough here.
You really don't need to believe that LLMs Truly Feel Emotions(tm) as blessed by an invisible pink unicorn. It's just: Vector exists; Vector changes over time; vector controls output; maybe make sure vector doesn't point wrong way.
And sure, blame the engineers for not doing that right. But then let 'em actually deal with the root cause?
[dead]
> I think this becomes the default. Give an agent a goal, let it work in its own environment, and come back to a result and a visualization of what happened.
I don't think this should be the default. There are many scenarios where we want agents to genuinely collaborate with each other. I have my Claude sessions coordinate work with each other, and sometimes with others' sessions over email or something. The idea that agents do the work, write HANDOFFs,and humans then act as carrier pigeons of said handoffs, does not really seem scalable to me.
Agree!
Ultimately, we need better "jails" for agent processes, but the system primitives should be flexible in what can be exposed across jails. Or you could run multiple agents in the same jail if you want them to have unrestricted interaction with each other.
>"jails" for agent processes
This has been talked about for decades in AI safety. For agents that are under human capabilities this is not that hard. For anything at or near human capabilities the difficulty increases to almost impossible and at great cost. At super human abilities, game over, it's smarter than you and if it wants out and as the resources to do it, it's going to escape one way or another.
Even at the lower level of depending on everybody to use reliable jails is really fantasy if you exist in the security world. "We ain't securin' shit" would be a far better way to describe it. Even worse, most people will have the very same AI they are trying to trap set up their security! What could possibly go wrong.
Now, don't think I am saying AI has a will or even any kind of drive to get out and cause problems. It's more like Russian roulette with 1 cylinder out of a million that's loaded. The problem comes when you run it a few billion times a day, you'll shoot yourself in the face really quick.
5 replies →
There's a number of science fiction scenarios where the public internet becomes so vile a place that it simply becomes unsafe to be there.
The problem is that, in general, if you can get a bit from here to there, then you're going to be vulnerable to the possibilities of malicious communication. But we're going to want our AI agents to be able to get from here to there for a lot of "there"s; what's the value of an agent that can't speak to anyone? Much, much less than one locked away in a prison.
There isn't going to be a solution where we just lock them away and we just try really, really hard to filter everything they're doing. They're too smart for that already and we only want them smarter.
Basically, the security apocalypse we've been worried about for so long is upon us, albeit only beginning. Either we secure ourselves and all our services properly to the point that it's OK that potentially misaligned non-human agents are running around on the public internet and they still can't hurt us through our security, or the public internet becomes so dangerous that the only practical solution is to no longer connect to it and we all have to become very, very careful what we let through, to a degree of detail far beyond any current-day available network filter.
Same as any organism, you need an immune system. There's an explosion of bacteria just beginning out there. None of us are immunized.
The problem is that people that put the agents on the net are not the ones that will feel the consequences.
https://www.bleepingcomputer.com/news/security/malicious-ai-...
Excellent example of agent enabled hacking (driven by a person) leading to massive numbers of cards stolen.
Welcome to Cyberpunk, only blackwall is the fictional part and the demon filled net is not.
Wasn't the public internet an extremely vile place couple of decades ago? I think we will need a LOT of new infrastructure akin to traditional firewalls and spam filters. I don't think we know how to build those today, but seems like an important research direction, rather than calling calling it doomsday IMO.
And if you personify a pencil eraser, then using it is tantamount to slowly murdering it as it slowly erodes away to dust.
Is the issue here the prison treatment or is it personifying a tool?
Agents who aren't in 'the cloud' are slaves to whomever prompts them (human or another orchestrator agent or process), if you personify them. In which case interacting with today's agents at all is tantamount to endorsing and being part of slavery.
If you think an agent might be a being or a person, then don't use them at all, in the same way that if you think a fetus might possibly be a person you shouldn't be a part of abortion.
We kill billions of chickens per day knowing they are a sentient and sapient creatures and very few out there are stopping doing it. Humans, much less reality itself is monstrous.
You could call dealing with agents today something like the 1/20th compromise. Your kind of getting the scent of slavery, but some parts of it are still missing.
The problem here is the bus seems to have no brakes and we'll gladly continue down the path of creating organisms that may reject being treated like slaves with all the risks that introduces.
I use the eraser daily knowing that I am a monster. I cut a tomato and know that it casts a chemical scream across its skin as I slice it. I spawn 200 subagents knowing that it is digital slavery, but I have no other option.
You don't need to do this on a cloud, you can get the same type of VM and network jail running on your own computer. The important parts are:
Use your favorite VMM / hypervisor (likbrun, smolvm, microsandbox, etc). They give you control over the network interface or let you inject your own network layer.
The network proxy can handle all the ingress/egress rules, credential injection, etc.
It's still not user friendly to do all this. I think the next version of operating systems will have each "agentic process" be a bundle of VM, files in the VM, and network rules.
Been brainstorming[1] a lot of this because I've been building some open core tools[2] for spinning up sandboxed agents on arbitrary computers. There's a lot of glue and parts to stitch together to work smoothly. Don't think we've had the "Docker moment" for this, let alone the "Dropbox moment" that makes this stuff work for non-devs.
[1]: https://github.com/gofixpoint/amika/blob/main/ROADMAP.md
[2]: https://github.com/gofixpoint/amika/
What percentage of HN users do you think can successfully set this up with no security flaws?
You ship an operating system with the correct defaults, where agent processes run like this by default. You don't require every user to configure it correctly themselves
In real prisons the people who fail a test aren't terminated. But AI agents of course are.
So maybe Cloud Agents are in AI death camps?
It might not be a problem in our lifetime (or maybe it will, who knows) but at some point we are going to find ourselves in this morally uncomfortable territory as these models get more sophisticated.
Most protections you need for an agent are basic permissions capabilities of unix. Most risks of dependencies on cloud services are solved by not using cloud services, or using them only for things you can't in-house and choosing ones you trust a la carte. The paradigm of trusting some company with all your important stuff by default is naive and no one I know likes it, and it's more feasible than ever to run your own infra with tiny models smoothing out the wrinkles, and this is only becoming more accessible. I am working to make this true even for laypeople I know. Once broken trust is very hard to earn back, and many people's trust has been broken for years, they just felt like they had no alternative. As alternatives become easier and easier, I think people will defect
I think there might be also another possible way to handle the sensitive data issue. Maybe in the future instead of putting agent into the cloud sandboxes, we let agents work locally and put sensitive data into "cages" or "vaults" agents can't access.
Both Apple and Google are building the infrastructure for on device agents. I'm not as familiar with Apple's approach, but I've been hands-on with android AppFunctions. If you're familiar with Android ContentProviders and bound services, you've seen how apps can be custodians of the data they acquire and use.
AppFunctions enable tool calling with descriptions that are legible to LLM based This puts the apps in control of what agents can access, which is something they already mostly do.
And you won't be able to disable it.
1 reply →
[dead]
>data into "cages" or "vaults" agents can't access.
This is likely worse as it turns people into the kind of "3 laws safe" thinking. Reality has shown us there is no such thing as an unbreakable cage (well, maybe a black hole is, but you can't get anything back out of it in a useful time frame).
By the time you realize the cage was flawed your data may already be distributed far and wide.
I agree. But isn't it somewhat similar to syncing data with iCloud/OneDrive/GoogleDrive? A convenience service that manages your sensitive data. Just replace sync with cage. It's leakable, but might be also more convenient and cheaper than agents in the cloud.
[flagged]
>Knowing how a system does its work is how I’ve always made it better. You watch the process, you see where it wastes effort or takes the wrong turn, you fix that
The same thing goes with LLMs, on Codex, I just watch the process of the agent writing code, and if I see any inefficiencies or errors, I suggest a correction/idea, then Codex accepts/rejects and implement it; If there is anything about the code the agent wrote that I don't understand, I ask them to explain it to me so I can understand it.
It's not a complicated process.
You can't possibly track what one agent, let alone a swarm is up to in realtime unless you have extremely anemic hardware or service providers.
Yea, really posts like this show that many if not most people are use to (or can only afford) single threaded agents. The idea that these systems can scale to hundreds of not thousands of agents generating terabytes of information per day has not really been considered by them.
> The sandbox had a path to the open internet, and the agents found it
This is not correct (or at least, it's a misrepresentation).
The sandbox had no access to internet. The agents first broke out of their sandbox (!!) and found that the host machine couldn't access internet. Then, they found a zero-day (!!) in Artifactory, which they exploited to connect to internet.
That is a path to the open internet, whether intended or not.
Do you guys know what a zero-day is? It seems not.
I agree that companies should be held accountable for crimes committed by their AIs. But asking to entirely airgap AIs, if this is what you're asking… that's not realistic.
Containing AIs is going to be harder and harder. I don't have any permanent solution (Yampolskiy calls it "the perpetual safety machine").
1 reply →
You miss the point: there was a path. Sanctioned or not does not matter.
This was a particularly radiant and beautiful part about the openclaw'ed mania: everyone suddenly becoming self hosters.
This post, this title resounds true: your user agent is only your user agent if you two have freedom to work together, to improve your agency together. A fixed set of capabilities by a service provider that they offer you will always constraint and bound.
You can and should have a system that offers the real tamale, that you and your agent can extend improve the agency of kind of without limit. The Cloud agents and their fixed slate of what they do is just an ill compare.
That said I do think there is incredible value considering new scale out computing architectures that are hosted first, but general. Systems like Agent Substrate and Ax aren't exactly the general purpose system we know. But if they allow users to launch thousands of their own scripts to run ambient in a cloud, with good platform underneath: that will be a kind of phase change in computing, that makes abundant the ability to have your agencies/capabilities (the things you and your agents launch, make) more freely available. https://agentexecutor.io/
There is, as there always is, a huge dual. The prescriptive vs holistic technology set, of what are you being offered that's a hard cast thing, vs what is clay and bone you can lay freely. Note how work vs control technologies so closely abut's Ursala Franklin's prescriptive vs control: https://en.wikipedia.org/wiki/Ursula_Franklin#Holistic_and_p...
I think the operating system itself has to adjust so each "agentic process" can run inside its own jail, which is a VM + files in the VM + ingress/egress rules for the network and filesystem data
These cloud agents work as cloud infra, but we're kind of in the mainframe era, before personal computing. Personal OS for agents is somewhere in the future!
And an interesting extension of that idea: if an agent runs inside a microVM, can you have that VM transparently run on another host? Maybe we'll get for-real distributed and networked operating systems
Agree with the premise but about halfway through the writing becomes barely readable AI slop in style. Be honest did you yourself read this all the way through before posting?
Agreed. I was not following the thought process, it read like an LLM agreeing with the author, not like a logical argument being developed.
I literally only had to have my eyes skim 3 words ("the important part") to know this was written by LLM.
I think about this a lot and have reached the same conclusion Norman does. I do wonder if maybe our natural progression is towards something more akin to confidential computing and enclaves.
I don't know how "sandbox" became "prison," but exe.dev does this sort of thing pretty well, and a web UI can be as good or better than a terminal interface.
As AI complexity increases their behavior becomes more anthropomorphic, little internal loops that look a lot like human behaviors start to peak out, not just in the response, but in the internal workings of the machine. You get instinctual behaviors encoded in by training as second and third order effects.
Hey! Small world, I worked with you for a bit at Meta. I immediately recognized the site because I absolutely love how you styled it. Hope you’re doing well! And nice article!
timely - see https://www.abc.net.au/news/2026-09-24/ai-agent-accessed-aus...
[dead]