Comment by probe
2 months ago
Do you think 3 is better than 1 & 2 as context gets larger? I think for smaller data sets its mostly fine no. It's an interesting bet by TM. End state does seem some form of continual learning (model weights update like dreaming)
No comments yet
Contribute on Hacker News ↗