Comment by applicative
6 hours ago
I had not heard of this comic masterpiece “Pfau et al. showed that a model whose chain of thought is just dots (“...”) can nonetheless … solve problems that are intractable for a model with an equivalent architecture but no chain of thought.”
The takeaway for that Pfau et al. paper is slightly more nuanced than that: It can only solve a subclass of problems without CoT, and that subclass can be equivalently solved with a larger model _without_ '...'
But arguably, a larger model will not need the chain of thought a smaller model does, which means simply by scaling we're already reducing CoT.
If the people who were relying on CoT are panicking now, they should've been panicking when perceptrons became multi-layer perceptrons.