← Back to context

Comment by glial

2 hours ago

'fast' means executing a policy, that is, a state-action mapping. A trained RL model does this.

'slow' means making one or several action-dependent forecasts, evaluating the expected value of the outcomes, and making a decision based on that.

Neither map exactly to the situation with LLMs, but very roughly, the first is analogous to trained classifiers and the second to reasoning models.

The analogy breaks down, since each instance of token being produced is an example of a policy execution (system 1), and reasoning is just stringing lots of these together. But there are those who argued, before LLMs, that system 2 is just "policy composition" anyway...

Wouldn't you need a classifier to even decide if it is system 1 or 2? How capable does this classier need to be?

  • I think of System 1 as a hash map. If you have a map, and see a new state/key whose action/value is not defined in the map, you have to go with System 2.