← Back to context

Comment by matusp

1 hour ago

> I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked.

He's not talking about decoders, he's talking about auto-regression. Before ChatGPT, the dominant paradigm was fine-tuning BERT-like models.

> Before ChatGPT there really wasn’t much of a concept of pre-training and post-training.

Again, people spend years just post-training BERTs in various ways.

GPT2 and 3 work via autoregression.

> people spend years just post-training BERTs in various ways

Yes, I was one of them. That’s not called “post-training” it’s called fine-tuning.