Comment by scruple
2 hours ago
I think that's equivocating on the word semantic. Word embeddings map concepts like "king - man + woman = queen" or cluster synonyms together in high-dimensional vector space and the ML literature very loosely calls this "semantic space." But a high-dimensional topology of token co-occurrences isn't semantics in the sense of computation or formal semantics. It's still "just" measuring distributional similarity. Some vector that represents "thread deadlock" lives near tokens like "mutex" and "race condition" and "starvation," but the model itself has no concept of concurrency and contention.
Claiming we "solved" statistical interpolation long ago just means curve-fitting and basic regressions on structured data. Transformers are a truly impressive achievement, scaling all of this to unstructured high-dimensional text topologies, but it's fundamentally the same math operations on statistical proximity.
Like how do we explain hallucinations here? Tokens that are hallucinated are semantically "close" in that vector space but they're completely false in reality. If LLMs operated in a true semantic space they wouldn't hallucinate CLI flags that don't exist.
> Some vector that represents "thread deadlock" lives near tokens like "mutex" and "race condition" and "starvation," but the model itself has no concept of concurrency and contention.
I propose that concepts of "concurrency" and "contention" are themselves vector in latent space. All concepts are. Recall that we're talking about a 10^4 - 10^5 dimensional space. You can fit in pretty much any conceivable association as some direction in there.
And try to zoom in on any concept you know. If you do, it should quickly become apparent that there's never any concept you can give a closed definition for. We can only define concepts, and we can only learn them, through generalizing from examples. Which is conceptually (pun not intended) regression - finding a vector along which examples live.