Comment by bigyabai
5 hours ago
You're kinda describing the MoE architecture; you can offload expert layers and stream them as-needed if the experts are small enough and the SSD is fast enough.
Dense LLMs typically perform better, but slow down much more than MoE models when you try offloading layers.
No comments yet
Contribute on Hacker News ↗