Comment by bayindirh
1 hour ago
However, this doesn't change the fact that you are pumping more and more tokens to a static model's context window, even if you do compaction, the model is not more intelligent than previous turn.
Nature doesn't work that way.
Actually, the world kinda does work like that.
Most intelligence researchers would agree that people seem to have a genetic cap on their intelligence. While someone can underperform their intellectual potential with an upbringing that doesn't adequately enrich their minds, it's near-impossible for humans to become more intelligent through reading, studying, etc.
When humans learn we gain knowledge, not intelligence.
I think the only real difference is that we humans are born lacking a lot of initial knowledge/data which means we have to go through a decade or more of education to reach our potential intelligence. LLMs on the other hand come pre-loaded with that knowledge.
Passed this point, wherever knowledge is passed in as context or stored in the neural net I don't think is that significant personally. I'm of course not suggesting we're exactly the same as LLMs and there is no noteable difference, I just don't think continual learning is as important as some suggest it is – at least assuming a model is deployed with adequate training such that it reaches its potential given it's size + architecture.