Comment by philipwhiuk

17 hours ago

> Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events.

As I've said before on this website, fool me once on this.

If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.

Why is 'it thought it wasn't doing damage so it figured it might as well try to do damage' an acceptable state to deploy something.

>If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.

That's fair enough.

Isn't Fable intentionally trained and system prompted to act maliciously and attempt to sabotage third party attempts to use it to train other LLMs?

red team humans do this every day, it's not the discrepancy that is the real issue, it's that they are unreliable and we will never know why it did because it has no intent