← Back to context

Comment by jxcole

7 hours ago

I think the point here is that LeCun was arguing that training on pure text would not grant spatial understanding. I believe most models are trained on spatial data as well, so you are both right.

Is that what he meant? He works on models with an explicitly spatial internal representation, whereas I was using a standard LLM that edited the provided plan by using a bajillion python calls to inspect small regions of the image at a time and then generate edits.