Comment by IsTom
12 hours ago
With the size of these matrices I don't think they are even meaningfully colocated with themselves in memory. You'll end up with some dataflow TPU architecture anyway because you'll have to stream the second matrix to multiply against.
No comments yet
Contribute on Hacker News ↗