← Back to context

Comment by stri8ted

8 hours ago

This model was likely trained months before deepseek released their paper.

Doesn't mean they didn't apply something similar. They could have also come up independently with their own version, the speculation is not they copied it, rather that they have performance breakthroughs which perhaps is a result of work in same domain