Comment by nomel
14 hours ago
What % of time, for a an average session, do you think is app overhead vs waiting for tokens? And there's your answer for why it's not a priority.
14 hours ago
What % of time, for a an average session, do you think is app overhead vs waiting for tokens? And there's your answer for why it's not a priority.
Jon Blow's response to this take was, "yes, which is why you have to work even harder to hide latency", instead of adding more on top.
From OpenAIs perspective, resources on your computer are free and wasting them is inconsequential
Kind of. Eventually it gets too slow even for the OpenAI engineers using it, and then they need to fix it.
[dead]