← Back to context

Comment by punnerud

3 hours ago

A trick to fit language models of this size it to remove the word to embeddings from the NN, and have a database to look up a dictionary of words and their embeddings. This way the LLL only need the "core" and you do the reverse for loopup again (kind of Text->RAG->LLM->RAG->Text). Have an example here: https://punnerud.github.io/pyspell/

The example have a limited language of around 1000 words, but make it possible to do Python (like) programming with LLM on an ESP32.

wow, that tailscale-in-browser demo is wild..!

re: language model - here, the embedding lookup, layers, logits all run on-chip and loops its own output back.

the only off-chip piece is the id to string table, i.e. the chip (once I get it!) will speak token ids and my laptop will print them as letters