Comment by SirHumphrey
19 hours ago
It’s a task much harder to RL and much more subjective. I don’t want to say we won’t get there, but let’s just say that LLMs could “write” well enough since gpt3.5 era and I don’t think the pleasantness of the prose improved dramatically since then.
And subjectively the explanation LLMs currently provide are usually horrible, horrible enough that I usually just instruct them to provide me human written literature I can read.
I mean, there's centuries' worth of mathematical prose to train on. But that's presumably already in the training data, so if it isn't good enough today, it might not get better fast enough to keep track with how fast they'll get better by training on formally verified math. But then again, the prose in Terry's conversation I linked above seemed pretty useful. But it's also a problem requiring famously little advanced mathematics.