Comment by joefourier
2 hours ago
I'm surprised you haven't heard about Natten, it's a few years old now: https://natten.org
It's the equivalent of sliding window attention, a common LLM optimisation. Global attention may be O(n2), but you can avoid that bottleneck by using local attention, even throughout the whole model. This could give you another efficiency boost, or enable you to process larger resolutions with linear instead of quadratic compute scaling.
No comments yet
Contribute on Hacker News ↗