← Back to context

Comment by hedgehog

1 day ago

No the weights are in the metal layers, they cannot be updated.

The base weights can't be updated but from what I recall it allows adding a low rank adapter to customize the model a little bit.

  • Yes that's what I've read. As far as I know the approach should transfer well to hybrid model architectures like modern Qwen and sizes like 27B by using multiple chips. LoRA-steered Qwen 27B at 10K+ tokens per second would be transformative for some workflows.