Comment by aogaili

2 days ago

Let me ask you, have you noticed that all the major labs are releasing models very similar in gains and performance? How do you explain that other than juicing the max out of the current architecture and processes.

They have been releasing similar gains/performance for years now... Pretty much never has one company been dominant for a long time (except the initial GPT3/4 release I suppose. That period took a while for others to catch up)

  • Yeah - that's why I think they are basically squeezing the scaling laws and the current architecture with incremental innovations.

    I expect they will continue optimizing and improving for the current use cases/benchmarks. But the core capabilities will stagnant, that's why they will be forced to slow down until another major breakthrough happens.

    I personally don't think LLMs have any understanding of the world, nor any imagination or even agency. I think it got very good as recognizing patterns and shapes in human thinking, language and knowledge, but that's pretty much it.