Comment by mmilunic
15 hours ago
Interesting paper by Thinking Machines where they solve this issue.
https://thinkingmachines.ai/blog/defeating-nondeterminism-in...
TLDR: It’s actually more about kernels changing with batch sizes, and you can solve it by making these kernels not depend on batch sizes. It took their inference time from 26s to 42s.
No comments yet
Contribute on Hacker News ↗