Comment by random__duck
1 day ago
Opened the RTL, looked at the floating point math, learned that apparently you don't need correct floating point operations for LLMs, closed the page.
1 day ago
Opened the RTL, looked at the floating point math, learned that apparently you don't need correct floating point operations for LLMs, closed the page.
Indeed there are some bugs/"non standard behavior" regarding very small or very big floating point values. All of those where proven harmless for LLM inference. Thanks for pointing out the bug. If you caught something outside of that, please point that out so I can fix it ;)
> All of those where proven harmless for LLM inference.
Could you share link(s) to those proof(s)?
Yup, the tests already compare to a CPU hugginsface implementation, but I'm working on making running and verifying the results of those tests easier and more available. Also, working on fixing those bugs and make the code more reliable now that the underling hardware has been saturated (memory bound)
Well, you don't. That's why 1.58-bit (ternary) quants are often used. But if that's what they're going for, no need to dress it up in a floating-point facade.