Comment by simianwords
10 hours ago
Asking earnestly, I don’t know what you mean by this reply. I know what system 1 and 2 is. But this has already been solved using same model.
10 hours ago
Asking earnestly, I don’t know what you mean by this reply. I know what system 1 and 2 is. But this has already been solved using same model.
> I know what system 1 and 2 is. But this has already been solved using same model.
No, it hasn't. Maybe you have a different definition of System 1 and System 2. I last read the book well over a decade ago (2011, maybe? 2012?), but System 1 and System 2 are different systems. IOW, System 2 is not a more computational version of System 1.
The argument you made implies that System 2 is just a more capable System 1, which is not what the book (nor this paper, AIUI) proposes.
In computery terms, System 1 runs in O(1) time, System 2 runs in O(log n) (or maybe just O(n)) time.
This means that any System 1 will run the input once through the heuristics, using the same computational power and taking the same time whether the input is 100 tokens or 1 million tokens, for quick but perhaps wrong decision (not "answer"). We don't have LLMs that do that. We have System 2 - run in O(log n) time and produce an answer.
System 1 is completely bereft of thought.
is this just a fancy way of saying that sometimes humans makes quick heuristic based decisions, and sometimes they think through things thoroughly?
I don't think there's any problem category that is strictly a quick heuristic decision or something that you will think through thoroughly always. I think it's more about how much time you have. If you don't have time you'll make a quick heuristic based decision. If you have more time you will think through it more. Imagine you're driving and suddenly like a branch falls on the road right in front of you. And you need to avoid it. Your brain will quickly use a heuristic based approach to avoid the branch. But on the other hand, if this branch was already there on the road and you saw it from far away, you will probably take a lot more time to think through and figure out which path you need to take.
The difference in humans is consciousness. System 1 is subconscious and fast, system 2 is conscious and slow.
System 1 can do multiple passes and all that, the speed is related to how aware are you of the computation occuring and system 2 must be consciously managed.
Example: You write a quick reply to a comment - thats system 1
You add two four digit numbers in your head - that will be system 2 unless you are very good, then it can be system 1 also
Why? Because memory allocations and addition computations are manually managed.
> but System 1 and System 2 are different systems.
They also share a lot of overlap in brain structures and they interplay while executing. A thought can start out Sys1 and quickly migrate to Sys2 as the pattern fails to match. Or a System 2 chain of thinking can be made from a bunch of smaller system 1 actions. Heck, in the middle of a system 2 thought you can plunge into system 1 system/actions. It's more of a who matches the pattern up with reality the fastest and acts on it.
> We don't have LLMs that do that. We have System 2
Eh. LLMs are system 1 thinkers by default. "Quick" response with no reflection is where LLMs started. It's later we added Chain of Thought and reflection layers and all kinds of other things like harnesses and agent training to make them act like system 2 thinkers. Of course we have other technologies being tested on LLMs these days like Dynamic Sparse Attention that likely match more of your thoughts on what system 1 thinking is too. Where the network doesn't have to parse the full context of the prompt and instead pattern matches with a much smaller percentage of the input giving responses back in ms versus seconds.
> The argument you made implies that System 2 is just a more capable System 1, which is not what the book (nor this paper, AIUI) proposes.
No, system 2 is the emergent capability to reason and increase the space of places to find the answer. Forget the paper's proposal, and look at the problem it is trying to solve. Ability to give quick answers, ability to give thought out answers, and the ability to know when to choose what. Adaptive reasoning does all three.
> This means that any System 1 will run the input once through the heuristics, using the same computational power and taking the same time whether the input is 100 tokens or 1 million tokens, for quick but perhaps wrong decision (not "answer").
No, I don't think we humans use o(1) to for understanding 1000 tokens or 2 tokens. I simply don't think that's the case. There's a new model called "Jev" and it is literally named System 1 (from the book) and even it is billed per input token.