← Back to context

Comment by ericd

4 hours ago

Nah, Taalas was putting the weights into silicon as a mask ROM. Their demo chip was hardwired to serve Llama 3.1 8B, and could never be updated. New models, even new versions without any architectural/size changes meant new tape outs.

But in exchange, you get insane speed and great energy efficiency. I could see it being a great approach for basic "good enough" models.

They may have had a little flexibility by supporting finetuning via LoRAs.

To qualify "insane speed" for anyone unaware, think a 10x improvement over even Cerebras. On the order of ~15,000 tokens/s. Not saying their approach scales well enough to keep pace with the various frontiers, but using their demo alone feels like a paradigm shift.