← Back to context

Comment by danielmarkbruce

1 day ago

Nathan Lambert wrote a good book recently, and he and his team wrote the paper below about Tulu 3 (Allen Institute). Both are good reads.

https://arxiv.org/pdf/2411.15124

Thank you for providing an arxiv!

An aside, I finally do appreciate single column format now, makes it easier to convert to epub.

  • When you are done with the section on RLVR, consider whether the model is predicting tokens, or making moves. There is a reason the word "policy" is used in RL.

    • Would still say it’s a token predictor, a fancy one though. I suppose we can agree to disagree.