Comment by HarHarVeryFunny
13 hours ago
It's a matter of degree not a black and while traceable/untraceable difference.
With a regular N-layer transformer you get a token every N-layers.
With a looped transformer there is no guarantee how often you get a token, unless you go out of your way to limit looping.
OpenAI's Jakub Pachocki says the "computational graph depth" (number of transformer layers passed though) for Astra is currently never more than 2x that of GPT-4, and does express concern that traceability will suffer if this is not controlled.
No comments yet
Contribute on Hacker News ↗