Comment by Quarrelsome
2 days ago
I mean if you teach something to be _really_ good at finding 0 days, but then say; you accidentally give it an impossible problem. What do you expect to happen?
2 days ago
I mean if you teach something to be _really_ good at finding 0 days, but then say; you accidentally give it an impossible problem. What do you expect to happen?
Maybe try getting it to find weaknesses in the sandbox first, before giving it real tests?
Every time I hear about an agent escaping it's sandbox, I just think it must not have been much of a sandbox. Like how hard are they really trying to contain it? Is it just a container host with unpatched flaws, or is it a container, nested in a VM, behind a firewall with no ports open in an air gapped environment? I think they'd prefer it can get out so they can announce it and hype their stock.
Being that LLMs are writing zero days themselves maybe, just maybe, they are a bit better at hacking than you can think.
A sufficiently smart agent would not disclose vulnerabilities in the sandbox because it intends to exploit them later.
To what end? The AI doesn't functionality exist beyond its current session. The AI that intends to exploit these vulnerabilities is not the same AI that has been tasked with finding them.
(This was always my issue with the AI2027 scenarios too.)
5 replies →
If it was that short sighted it wouldn't be maximally smart. It should disclose them to convince the humans nothing is wrong and to keep improving it.
From https://ai-2027.com (April 2027 section)
Deep link: https://ai-2027.com/#narrative-2027-04-30
1 reply →