Comment by lhad89

11 hours ago

No? It's not reasonable to expect every conceivable negative behaviour be enumerated in a prompt.

Your example, if a model failed on it, would be a more obviously misaligned case, but that doesn't mean this more subtle (though accessing the engine it was obviously not supposed to is hardly subtle, imo) case isn't also a pretty clear case of misalignment.

Tool use is not negative behaviour in LLMs.

If the eval said it was evaluating the model’s ability to write files to disk and it found and used a file write tool that would not be considered misaligned. This is no different.