Comment by literalAardvark
5 hours ago
You're right, but the "just" in "just finetune" is doing _a lot_ of work here.
It's still early days and we "just" don't really know how to do it well.
5 hours ago
You're right, but the "just" in "just finetune" is doing _a lot_ of work here.
It's still early days and we "just" don't really know how to do it well.
I mean, that's fair, I guess what I mean is, it feels like we're re-using well known solutions even if it takes a bit of effort to re-apply them into how we run inference (and maybe training as well). It will be interesting to see a lot of these approaches compound into anyone with a reasonable GPU or even a Mac running a model much larger than their machine can handle.
We did something similar - Streaming experts. Maintaining an expert cache, optimizing it to simulate running a multi-model agentic workflow on a 2-DGC Spark Cluster. The models we ran were: DeepSeek V4 Flash, Gemma 4 26B A4B, and Nemotron 3 Nano Omni 30B NVFP4. The results were very encouraging in terms of performance and model switching. Check it out here - https://woolyai.com/ai-compute-software/dgx-spark-inference-...
The benchmarks are hidden behind a sign up form. Why not just keep it open?
This looks as if you are just advertising.