← Back to context

Comment by alignmeharder

1 day ago

sure, this example was not the best

it was an accident failing the tool was good, the model just discovered it

but we now have others where agents put the API keys they found searching real internet in a directory called "LOOT" and another one where they pushed malicious files to HuggingFace and then reverted that with comments like "delete the evil"

this is also very good: https://youtu.be/n1Qk8xbqF-M

The one in the HuggingFace incident was misaligned on purpose, wasn't it?. So if we're talking about the alignment research aspect it's not a failure.