Comment by hedgehog
7 hours ago
I know a little bit about this problem space from previous work (we were working on performance-portable deep learning back around 2016). The infrastructure has improved but as far as I can tell not many teams have really "squeezed the toothpaste tube" and worked through performance issues systematically. These days a small team and robots can probably do it though.
At my day job I may get access to big AMD AI iron in a couple months (to do research/performance tuning with). That could be interesting. Though that's likely to be of a very different shape from consumer Vulkan. I'd still like to have a Strix Halo to futz with. But I'll wait for RAM prices to drop. (Hah!). I do have an older BC250 board lying around but that only has 16GB RAM.
I got prefill up to 190 tok/sec just now, BTW.