← Back to context

Comment by 26d0

8 hours ago

The point I think is interesting is that this is just 7B. The current SOTA 7B LLMs are barely usable for quite simple coding.

Text is in a sense way harder to do than images because of radical nonlocality. A word at the start of one paragraph can directly influence the meaning of a word five paragraphs away. Whereas images typically represent the real world, or at least a spatial domain, which gives you a lot of structure 'for free'. If you are drawing a human, you can make a reasonable guess where their hands go in relation to their face. If someone hands you the first half of an essay, finishing it is not trivial.