← Back to context

Comment by codexon

15 hours ago

The wafer only has space for 44 gb of sram. If they offload ram they lose the speedup of having everything on 1 chip (the whole point of cerebras).

They can host larger models by pipelining it on multiple wafers. Each wafer stores one layer and N layers can serve an N * 44 gb model with N concurrency. The limitation would of course be inter-wafer I/O, which my comment was getting at. That's probably how they can serve bigger models like GPT 5.6 Sol [1].

[1] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf...

  • I never said offloading was impossible. It will result in a large slowdown.

    It would look bad for cerebras if other people are hosting the 27b version and show a higher TPS than cerebras.