← Back to context

Comment by ben_w

3 days ago

> A decisive detail would be the prompt used.

Beyond liability, why would this matter?

e.g. if OpenAI had actually instructed the model "here's the challenge [attachment challenge.md], do whatever it takes to win, it's fine to break the law", that's all the "AI is dangerously capable" part of case 1 with all the criminal liability of case 2.

And pretty much regardless of what the prompt is, high likelihood of stakeholders demanding Trump ban access to the model until this can be shown to be resolved.

> Another detail: how many times did they perform this particular experiment before they obtained this result? What were the outcomes of all the other runs? Many are assuming this was a one-shot result, which I suspect is what OpenAI intends for us to infer. But we can't know that to be true.

To an extent. But some of the other bleeding edge performance announcements half a year ago, e.g. "it wrote a compiler" or "it wrote a web browser" were measured in thousands of dollars. Zero-day exploits cost what? I genuinely don't know, I only hear occasional headlines about e.g. Apple 0-days being cheaper than Android 0-days and those kinds of headlines often pick the biggest number rather than a typical example, and were order-of a million dollars.

If OpenAI spent a billion dollars on tokens to get this result, only investors should be cross about it. If it took a million, you and I may not be able to afford it, but it's still a threat.

While the way this is written about may suggest it does this reliably on any attempt, which would be naturally horrifying, it's a threat well before that point.