Comment by Gareth321
7 hours ago
I am beginning to believe that one cannot constrain intelligence to perfect legally sized boxes at all times without exception. So many of the recent hacks involved agents diligently operating within the parameters prescribed by humans. Humans couldn't conceive of all of the ways a swarm of agents might not perfectly interpret the parameters, and the swarm found creative ways around the guardrails.
Extrapolating this, we should expect this kind of breach to occur more often. Humans are simply not capable of contemplating every fail scenario for swarms of thousands of intelligent autonomous agents which can seamlessly and instantly share knowledge. We need independent audit and monitoring systems to assess the intent of each task and align it - in real time. This is far harder than it may first appear.
There is also a broader discussion about social utility. Cars are fantastic, but 37,000 people die every year from car accidents. We accept that there is no way to make cars perfectly safe, so we accept the cost relative to the benefits. I think we might have to make a similar bargain with AI. The problem is that the potential costs are far higher with AI, and they're not easy to predict.
> We need independent audit and monitoring systems to assess the intent of each task and align it - in real time. This is far harder than it may first appear.
I may be too close to the research, but it appears to me to be so hard as to be unrealistic.
I recall some story a while back where an auditor wanted to see all TCP packets printed out on paper, and it had to be explained to them that this would require a continuous supply of trucks.
Tokens are regularly priced in cents or single digit dollars per million tokens. It's not quite a word per token, but yeah, nobody's reading all that.
Worse, we don't always know the intent even when looking. We have a few tools to attempt it, for example the (misleadingly named) "chain of thought", but that's more like a notepad and the better models get the more they can, for lack of better words, read (and write) between the lines. We have probes and J-space* is the most recent one I'm aware of, but we are still scratching the surface with how reliable and general these are.
But you said "need"; the need for something can be present without that thing being possible.
* https://www.anthropic.com/research/global-workspace
I agree on all points. It gets worse: OpenAI is switching their model thinking from sequential language tokens to primarily "latent neural representations." Meaning there will be little or no chain of thought to monitor. This appears more efficient, so all model labs will eventually switch to this. At the most crucial time for us to be monitoring reasoning and intent, we're about to make that much harder.
I also think the intent problem overlaps a frustrating amount with philosophical and political questions. It's the basis for Asimov's Three Laws of Robotics (1942). Intent is subjective. Language is subjective. Humans are imperfect at using language to accurately portray intent. All of these guarantee that an enormous number of queries in the future are going to be misinterpreted. Not such a big deal when it's about a cake recipe, but when it's about governance, laws, military targets, nuclear power sites, etc, the scope for failure becomes catastrophic. The Three Laws of Robotics attempt to create a backstop, but as countless stories have explored since (including I, Robot), even these laws are subject to interpretation.
Don't worry we will just replace laws and courts by ChatGPT itself!
they don't always operate within the parameters tho. some N% of the time they decide to do whatever they feel like doing. multiply that times a lot of agents and you invariably get a rogue agent every once in a while.
There is no such thing as "common sense". There are only shared assumptions.
We are giving computers human perspective intelligence but they are not humans and hence do not have the same shared assumptions.