Comment by malwrar
21 hours ago
What about when the open models get to the point that you don’t need these fancy GPUs? It seems inevitable this will happen over time, we see cracks right now within even just GPT w/ expert streaming and flash attention. ASICs are coming soon too.
Nvidia already have a plan in place for fast inference. They acquihire Groq last year