Comment by l33tman
3 years ago
I know what you're saying, and for sure existing models can be difficult to force into the really weird corners of the distributions (or go outside the distributions). The text interfaces are partially to blame for this though, you can take the images into Gimp and do some crude naive modifications and bring them back and the model will usually happily complete the "out-of-distribution" ideas. The Stable Diffusion toolboxes have evolved far away from the original simple text2image interfaces that midjourney and dalle use.
The models will generalize (because that's the most efficient way of storing concepts) and you can make an argument that that means they understand a concept. Claiming "it's not learning concepts, only statistical probabilities" trivialises what a modern neural network with billions of parameters and dozens of layers is capable of doing. If a model learns how to put a concept like line width 5, 10 and 15 pixels into a continuous internal latent property, you can probably go outside this at inference at least partially.
I would argue that improving this is at this point more about engineering and less about some underlying unreconcilable differences. At the very least we learn a lot about what exactly generalization and learning means.
No comments yet
Contribute on Hacker News ↗