Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model.
Maybe still worth it if their "64% cheaper" figure holds.
I guess I don't. Does post-training from another (larger) model not fall under the umbrella of distillation? I'd imagine it leads to the same spiky-ness issues...?
I presume post training is significantly easier than the distillation/training the top Chinese labs are doing.
I wonder if, similar to the American labs, they'll become stingy with their weights once they start getting immediately undercut by a wave of slightly better derived models.
Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model.
Maybe still worth it if their "64% cheaper" figure holds.
I don't think you know what distill means
I guess I don't. Does post-training from another (larger) model not fall under the umbrella of distillation? I'd imagine it leads to the same spiky-ness issues...?
3 replies →
With the Devin subscription even at the 20$ plan, they offered unlimited SWE 1.7 usage. Wondering if they do the same for SWE 2.
SWE-2 is free for all subscribers on the CLI to try out for the next month :)
1 reply →
I presume post training is significantly easier than the distillation/training the top Chinese labs are doing.
I wonder if, similar to the American labs, they'll become stingy with their weights once they start getting immediately undercut by a wave of slightly better derived models.