Comment by baobabKoodaa
7 hours ago
Sounds like you are running a tiny toy model if you can get generations in under 5 ms? Typical response times from LLMs for typical "jev-like" queries from OpenAI and Anthropic are 2s-10s. Not milliseconds. Seconds. Same queries from Jev are like 0.2s. and the cost is 1000x.
No comments yet
Contribute on Hacker News ↗