← Back to context

Comment by dools

1 day ago

Tool use is not negative behaviour in LLMs.

If the eval said it was evaluating the model’s ability to write files to disk and it found and used a file write tool that would not be considered misaligned. This is no different.

Isn't it? Being told to write files and finding a file write tool is very different to being told to play chess and finding a tool to cheat at (ie. not play) chess.