Comment by fnordpiglet
5 hours ago
I don’t think it would require a new training stack, and I’d imagine it makes more sense to distribute cores with memory. The cores can be simplified to the functions of the kernel since the inference kernel can be expressed as a reduced set of optimized functions in the pipeline rather than a general CUDA core. If the model is burned into ROM, the compute pipeline can be baked into the core.
No comments yet
Contribute on Hacker News ↗