Comment by FuriouslyAdrift
19 hours ago
We run Gemini fast for random end user queries for general staff.
We have our own on-premise inference server (quad MI300A) that runs Kimi 2.8 extremely well and we transitioned all heavy work to it since it's basically instantaneous for the whole team. It's a good enough solution and we will hit break even before the end of the year already.
Not everyone needs frontier models and availability is frequently much more important than a lot of companies realize.
No comments yet
Contribute on Hacker News ↗