Comment by r_lee
4 hours ago
if it's agentic stuff, they likely aren't hammering a core constantly and they will maybe sit idle quite often between model requests, so it makes sense. I do wonder how much memory they allocate to each one though.
it's just very efficient use of shared cores that is required to make these kinds of workloads cost efficient
No comments yet
Contribute on Hacker News ↗