Comment by HarHarVeryFunny
2 hours ago
If you understand what they have achieved here, then the notion that they are bottle-necked on training data is absurd.
I wonder how you imagine that China built their own space station? Reliant on using American made duct tape, perhaps?
Do you realize how reasoning models are being trained nowadays? You design/build simulation environments to run agents in, with the environment providing the RLVR "verification" scoring. So why won't Ziphu use GLM to build their own RL training environments? Do you think they are not doing this?
No comments yet
Contribute on Hacker News ↗