← Back to context

Comment by walrus01

8 hours ago

some people made a 'caveman' speak qwen as a joke

https://huggingface.co/ProCreations/grug-27b

It's not exactly a joke, it does reduce the amount of tokens. However, it does not improve performance (fine tunes are finnecky things, hard to get one right).

  • Personally the only 'enthusiast' modified qwen 3.6 27b or 3.6 35b-a3b I've found useful are the ones that have been run through heretic and adversarial data sets for innocent/dangerous prompts, to produce uncensored LLMs. They have some niche non-coding uses for things that a commercial LLM will never talk about.

    https://github.com/p-e-w/heretic

    • I think those are mostly vapor that runs on the small culture of "models should not be censored" thing. But from my experience, they unlock nothing meaningful.

      Fine-tuning is great for really small models on specific applications, but it's not something that can essentially improve a more generic model.

      That said, there seems to be a fine line in quantization+finetuning that could recover performance. It's just hard to get a hold of it (I feel it in some models, but it's hard to say yet; lots of small labs working on this RN).

      3 replies →