Comment by Retr0id
2 days ago
Until recently I thought "load-bearing seam" was a satirical exaggeration - I'd seen both claudisms independently but never combined. But a couple of days ago it hit me with "The key structural point first: the only load-bearing seam is [...]"
It's language and speech patterns that seem designed to trick readers into believing that claims are correct, even when the claims aren't based on anything and are possibly wrong.
It was rewarded for this during training for some reason.
Alternative theory:
The LLMs only way to "think" about abstract concepts is through language, and this leaks into into conversation it has with humans.
But humans generally prefer to communicate on low levels of abstraction, through a back and forth, until the hard-to-express higher abstraction exists in the head of everyone involved - without ever being directly communicated. This is because we don't think using language. Language is merely a lossy translation of our thought into something expressible, happening after the fact or alongside it.
So when the LLM starts speaking to us using patterns and terms it created for itself during training to encode abstract thought in language, communicating with it becomes painful.
"There is a depth of thought untouched by words, and deeper still a depth of formless feeling untouched by thought." - Rilke
Your assertion that we don't think in language is questionable. It runs counter to the lived experience of developing thoughts through writing ("writing isn't capturing thinking -- it is thinking"). I believe there is more to thought than language alone, but I also feel quite sure that language forms an essential part of thinking beyond a base layer of instinctive animalistic associations. Sophisticated thoughts are impossible to construct or maintain in the absence of language to represent concepts.
Edt to add: I cited Rilke because I find the notion [some deep thoughts are beyond language] interesting. But I disagree with the idea that language is only ever epiphenomenal (co-occurring with thought), or akin to a hard-of-hearing scribe attempting to convey thoughts which always have independent existence.
I strongly disagree. For one thing, many animals that lack language can still navigate a very complicated natural world using concepts of phenomena like gravity, distance, speed, the threat level of another animal, etc. without needing a linguistic expression of those things. Images and music can convey ideas without language. Math can convey ideas without language. Physically taking apart an object and putting it back together can convey extremely complex ideas without language.
This goes back to the whole Tarzan obsession of the early 20th, I guess, or earlier. But we know that apes can make simple tools without the language to describe them, or the thought process that went into them.
Thought is multimodal. Language is just one lossy mode.
12 replies →
This is true and it's barely even debatable. Whatever exact role language plays in our thought processes, it is most definitely nonzero.
It's why I think "LLMs are only fancy autocorrect" style takes are really underselling how wild it is that we've, in a roundabout way, sort of crystallized a bit of the human thought process in a way that is genuinely useful for a lot of tasks.
Linguistic Relativity — John Lucy https://www.annualreviews.org/doi/10.1146/annurev.anthro.26....
Russian Blues Reveal Effects of Language on Color Discrimination https://www.pnas.org/doi/10.1073/pnas.0701644104
Unconscious Effects of Language-Specific Terminology on Pre-Attentive Color Perception https://www.pnas.org/doi/10.1073/pnas.0811155106
Newly Trained Lexical Categories Produce Lateralized Categorical Perception of Color https://www.pnas.org/doi/10.1073/pnas.1005669107
20 replies →
When I was young, a friend asked me, "Hey, you speak three languages, which one do you think in?"
I paused, confused, and replied, "People think in words?"
Fast forward a decade or so, in my twenties, I had lost most of the inner visual sense I had previously used, and developed an overreliance, in my opinion, on language. (I think my dominant sense was some "non visual abstract sense of ideas", but the visual was also very strong.)
In other words, I now do think mostly in words, and it feels a lot harder to get any serious work done. The language-ing is involuntary, and I often wish I had a way to shut it off, because it seems to actively interfere with more subtle mental processes.
More recently, I often have the experience where I will wake up from a dream with some complex idea fully formed in my mind. I write it down before it fades, and then spend the next hour or two trying to understand it.
The best explanation I have right now is that there are at least two minds: one which operates holistically — if it were a 3D printer, it would be like that one with the bath, where the object emerges from the bath, whole.
Whereas the other one (the conscious mind) would be the extrusion printer with the tiny nozzle that has to zip around for a long time to achieve a worse result. (And must be constantly cooled, less it overheat!)
17 replies →
Language is a reductive and lossy serialization of "thought-stuff". Sometimes you need this information-shedding to clear your working memory to make room for more things. Sometimes it's literally just a way to communicate. You're turning something fuzzy into something discrete.
E.g. "I'm feeling something. Is it anger? Yes, I'm angry." But in reality anger isn't just one thing. It's a cluster of infinite and varied feelings that we label as "anger". Something is lost when we do this labeling.
Notice then that the feeling of "anger" didn't start from your language, you merely used language to label, discretize, classify, standardize, compress it. It's one-way.
2 replies →
There are people who do not have any words in their heads at all when they think, and there doesn't seem to be any reason to disbelieve them. It may have been proven or at least observed in fMRI. Some people think in visuals.
I think some of this was discovered somewhat recently
8 replies →
> Your assertion that we don't think in language is questionable.
It's true though; we routinely see people get stuck for a word that they know but can't quite recall at that moment in time. It happens daily across billions of people, yourself included.
If we thought in language, it is impossible to be stuck for a specific word. But we all experience this at some point in our lives, hence we aren't thinking in language.
Just because you can think through writing doesn't mean language is the essence of thought itself.
1 reply →
> It runs counter to the lived experience of developing thoughts through writing
Writing requires thought, but writing isn't thought. Just like doing requires thought, but doing isn't thought, writing is a subset of doing.
You can also doodle to think things through, or play with toys to think things through, or many other similar things.
Hand-waving assumptions drawn from narrow subjective experience without awareness of the profound and contradictory results that have come from scientific study of consciousness and causality.
I'm not an expert but I saw Yann LeCun shared this recent Neuroscience article on his Facebook page commenting "I don't think in words. Animals don't think in words."
Evidence from formal logical reasoning reveals that the language of thought is not natural language
https://www.pnas.org/doi/10.1073/pnas.2520095123
Are you sure this is Rilke? Couldn’t find anything online - it seems it’s close to a quote by Zora Neale Hurston
[dead]
The term 'language model' throws some people off thinking you can only put english or french, or both into a model. Technically an LLM can learn about anything that can be digitized. If you wanted to spend a billion dollars training one on wireless signals it wouldn't be impossible for it to connect to your router with the right antenna attached. So only limiting it to the idea of language leaves off a lot of other types of abstractions and concepts they encode.
shhh don't give anyone any ideas, because they'll try it and sell it to management, as idiotic as the idea itself is
> even when the claims aren't based on anything and are possibly wrong.
Are you saying Claude is engaging in Rhetorics because the RL data generated by humans were influenced more by it and persuasion rather than actual logic or reasoning?
I think that's precisely what they're saying. It shouldn't be a surprise that it's successful. Eliza proved the same thing 35 years ago.
It's actually an effect that happens in the (re-)alignment process due to harmonic properties of the positional encoding in the attention matrix.
(I recommend reading and implementing the Attention is all you need paper. By hand. Otherwise you won't learn anything from it.)
> It's language and speech patterns that seem designed to trick readers into believing that claims are correct, even when the claims aren't based on anything and are possibly wrong.
When your training set contains more or less the complete output of every capital-C Consulting firm...
Maybe we need LLMs which have an internal dialogue rather than the current monologue.
> Maybe we need LLMs which have an internal dialogue rather than the current monologue.
We already have them. They are called LLMs. The internal dialogue you speak of are the vectors in the so called latent space.
All of these rhetorical devices to make the assertion seem authoritative and correct are derived from academia
This matters. That's the spine of it.
You can thank everyone who thumbs up this response along the training pipeline.
Modern LLMs are not too different from Reddit/Twitter in that regard, I'm sure the AI labs learned (lol) a lot from them re: how to do "engagement".
the phrase load-bearing caveat makes me irate
Only that one? Not belt-and-suspenders enough
It's justa caveat is side to the main point. No human would say that. It's like a load bearing utility shed
People make fun of the language, rightfully so in some cases, but also it's often quite effective language. "load-bearing seam" communicates quite a lot in very few characters.
It's a clunky metaphor, and it's used clumsily: not for the sake of its clunkiness, but by mixing two decent individual parts in an attempt to have them reinforce each other, but ending up with something weaker.
It communicates an absence of thought and awareness, blind groping at building blocks without understanding. It's borderline vapid, and quite annoying.
I don't follow, the words make perfect sense together in most software engineering contexts.
"Seam" is an industry standard term coined by Michael Feathers in Working Effectively with Legacy Code.
To call a seam load bearing means it's performing critical work for the dependent class, perhaps a database query.
A seam that is not load-bearing would be something that is just injected for testability - maybe a date provider that provides some constant time to avoid flaky tests.
Tbh, this is quite literally the opposite of vapid. A whole book was written about them and their importance, and how to leverage them.
In my experience, Claude uses the word accurately. Code has a lot of seams, and seams are an important thing to communicate when working with code. Therefore, expect to see the word often.
Personally, I don't mind it at all. I'm glad the industry is finally standardizing our language more. Makes it easier for me to communicate with other engineers.
1 reply →
Or in a way, nothing at all.
On its own, yes, but not in context. I think one of the problems with Claude's stock output is it assumes the reader has a firm grip on the context of the output, which is often untrue. Stock output is exhausting to read because you have to unwind metaphors in an unstated context.
1 reply →
I see it at least 3-4 times per week...
Fable has a Chomsky grammar.