Comment by BoorishBears

1 hour ago

Expecting a strong zero-shot performer to perform worse in a low data regime?

That only makes sense if you try to rope in data previously used to establish the model's priors, but that wouldn't make sense in this context. That same additional data is what enables things like...

> use generalized models to generate ad hoc specialized classifiers.