Comment by porphyra
11 hours ago
They can host larger models by pipelining it on multiple wafers. Each wafer stores one layer and N layers can serve an N * 44 gb model with N concurrency. The limitation would of course be inter-wafer I/O, which my comment was getting at. That's probably how they can serve bigger models like GPT 5.6 Sol [1].
[1] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf...
I never said offloading was impossible. It will result in a large slowdown.
It would look bad for cerebras if other people are hosting the 27b version and show a higher TPS than cerebras.