← Back to context

Comment by riedel

6 hours ago

Can someone explain to me how such self-awareness can be forced into the model. I mean I guess the pre training data could contain all sorts of stuff. How reliable are those hacks. I know that a lot of open weight models answer that they are Claude in the absence of a system prompt. I find destillation not that much of a plausible explanation as typically claude would probably not mention that it is Claude all the time. I find it rather plausible that a foreig. system prompt made it into pre-training. But again: I have no clue how much care is given by models to leave traces for destillation (for closed weights) or post training (for open weights).