← Back to context

Comment by cfiggers

2 hours ago

Consider this "strong" premise:

    So-branded "AI" models (that is, big bags of numbers and the linear algebra needed to exercise them) can be inherently or autonomously dangerous at a scale that threatens the long-term survival of the human race.

Consider also this "weaker" form of the same argument:

    So-branded "AI" models can be inherently or autonomously dangerous, if not at a scale that threatens the long-term survival of the human race, then at least at a level that significantly threatens, inconveniences, or harms humankind to a noteworthy extent.

Here's a few hard truths that no one will like to hear:

1. Those who accept and advocate for either or both of these premises do so based on ideology, and not based on the scientific method.

This is not a controversial statement. Advocates of the "AI danger" premise (whether the strong or weak forms above) freely and gladly, even insistently, note that the conclusive proof of what they're describing has never been observed, much less observed enough for any kind of experimentation loop to have run for any meaningful number of iterations. They are generally fine with accepting the premises without conclusive proof, because, they theorize, the instant that a conclusive proof event occurs and we observe it for the first time, humanity goes extinct shortly after. So if either premise is true, the consequences of them being true can only be averted by accepting the possibility / probability / certainty of them being true and acting to prevent their consequences without waiting for conclusive proof.

2. Statement 1, which asserts a brute fact and not my or anyone else's opinion, is not a commentary on whether either premise is in fact true or not. It is only a commentary on the epistemics of those who broadly accept them.

3. Since the conclusive proof is not available, advocates of the "AI danger" premise argue publicly for their view based on what they acknowledge (again, freely and gladly) to be inconclusive proofs.

Popular in this set of inconclusive proofs is the Hugging Face hack, which allegedly demonstrates that so-branded "AI" technologies are capable of becoming inherently or autonomously dangerous at a scale that threatens humanity to either the greater or lesser degree above (this is a different assertion than the one that claims they're already that inherently or autonomously dangerous today).

4. While we have plenty of facts about these inconclusive proof events, the facts most critical to their interpretation as supporting the premises above come from biased/non-neutral sources.

The chain of events that transpired in the Hugging Face attack are publicly documented perhaps better than any other cybersecurity event in history. But the details that make this event as a whole either compelling or not compelling in support of the "AI danger" premise (either strong or weak form) come from within OpenAI itself, and are prone just as much to selective omission as direct manipulation. These are details like:

- What exactly were the models that exhibited this behavior, and what were they pre- and post-trained on?

- Were the agents pushed at all by human intention toward creating a public spectacle, the way they subsequently did?

- Why were so many agents run on this task in parallel and what did OpenAI expect the return on investment of so much compute world be?

- Was the weak sandboxing known about prior to the attack? Was it purposefully either arranged that way or noticed and not fixed?

Reliable answers to any of these questions world completely change the salience of the Hugging Face attack to the "AI danger" premise, and on every one of them we have to take OpenAI's word for it (or else be content when they choose not to address it one way or the other).

So in brief: we don't know anything for sure, what we'd like to think we know comes from unreliable sources, the loudest voices advocating for the most sweeping change are ideologically driven, and everybody with access to ground truth has maximal incentive to blur, bend, or break the conveyance of that truth to the public. So what do we even expect for our own ability to discern reality from fiction on this topic? We should expect little, if any ability at all.