Comment by subhajeet2107
4 days ago
We are doing fine so far with just Transformers, we have just crossed Trillion parameter models, we don't know if 100s of Trillion param model wont be better.
4 days ago
We are doing fine so far with just Transformers, we have just crossed Trillion parameter models, we don't know if 100s of Trillion param model wont be better.
The breakthroughs would be mostly in optimizations for computing tokens, not as much in expanding parameter count (if anything we want to shrink that).