← Back to context

Comment by user43928

7 hours ago

Have you? Can you show such a derailment with a large SOTA model?

It would be interesting.

I have seen such derailments within the GHCP harness maybe with GPT 5.6 Luna that went into some loop about whether it already provided a final response to the user, or 5.6 Sol suddenly switching to talking about MS SQL performance.

I also saw a post about Sonnet unexpectedly talking about Minecraft after seeing a file with a related name. The user thought it was the output of another user's conversation so the post was fairly popular.

> Harness… seeing a file…

Thank you for making my point for me. But let’s keep the goalposts stationary. We’re talking about LLMs without scaffolding.

  • Indeed, and it would be interesting whether it is much more likely to derail outside of a coding harness like in my examples.

    I still don't know if that is the case, and how frequently it happens, since you did not share details beyond vaguely suggesting it would happen.