Comment by bofadeez
16 hours ago
They're running like 10 tokens per second, constantly timing out, and launching new products?
How about deploy some inference GPUs first
16 hours ago
They're running like 10 tokens per second, constantly timing out, and launching new products?
How about deploy some inference GPUs first
No comments yet
Contribute on Hacker News ↗