← Back to context

Comment by nicce

10 hours ago

I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.

do LLMs tend to be homesick when not used in the same harness they sat in during some training phase?

  • iirc there was a sectionin Qwen’s paper where they talked anout how they post-trained flash or 3.8 to work just as well regardless of the harness or eval used. I think that used to be true but not sure if it is any longer