Comment by tidewave
13 hours ago
Congrats on the release!
Finetuning a language model for decision classification (with probabilities) is already well-understood. What specifically changes in the training objective with RLCD? Are its benefits isolated from Jev’s new architecture/parallelism?
No comments yet
Contribute on Hacker News ↗