Comment by bonplan23
5 days ago
Just to be clear: It was well known that you can reach such scores with small models and without an LLM if you train on the task. The author highlights those models himself - e.g. HRM/TRM.
The novelty is more that it works with such a plain transformer and low compute price.
No comments yet
Contribute on Hacker News ↗