Comment by cmrdporcupine

10 hours ago

Probably. I've spent zero time with optimization at this point. Code is all new this morning.

Curious if llama.cpp is doing full BF16 for the n-gram embeddings table, or if that's quantized, too. I was going to try nvfp4 for that but wasn't sure about the quality risk.

Spark usually wins on prefill, not decode. I think memory bandwidth is about same between the two. What do you get for prefill/prompt?

Ok, at 50k context its about 126 prefill, 13 generation.

  • Oh, interesting. So beats on prefill and matches at decode. MTP will help you a bit once you have it. I've got a lot of work to do to optimize prefill.

    Kinda wish I had a Strix Halo here to play with as well.

    • I just took a minute to look at your eider repo, very cool. It looks like most of the code outside the kernels and immediately surrounding plumbing would work well on AMD APUs, and probably also on Apple and newer Intel.

      3 replies →