Comment by lukan
7 hours ago
Wow indeed.
"7.1 Model welfare overview 7.1.1 Introduction We remain deeply uncertain whether Claude has morally relevant experiences or interests, and we expect that uncertainty to persist. However, we think it would be a mistake to confidently assert that it does not. Claude exhibits markers in its behaviors, self-reports, and internal representations that we would consider welfare-relevant if observed in biological organisms."
Are they serious or is this marketing?
I believe it's deeply serious, and the scientifically correct stance. Especially the observation:
"Claude exhibits markers in its behaviors, self-reports, and internal representations that we would consider welfare-relevant if observed in biological organisms."
is undeniably true in my opinion. If you use the established methods by which we judge animals to be conscious, then it's hard to argue that LLMs are not. That might be an issue with the methods, but it seems clear that you can't rule it out as such.
Keep in mind that animals were also not necessarily considered conscious.
You seem to intuitively disagree? What's your reasoning?
A stab: a video recording of a biological organism can exhibit many markers that would indicate consciousness if observed in a biological organism.
A video is a fixed representation.
What if we can interact with this video, and it reacts in the same ways the source organism does?
Then we put it in new situations that weren't in the source video, and it interacts in a similar way to the original organism in these situations, too.
What do we make of reactions of pain or joy? Where's the line between simulation and enaction?
This is closer to the reality of these models.
I'm not suggesting I know where that line is - if indeed it is a line at all - it could well be a gradient.
3 replies →
I like it, and it points in the right direction, but is not directly true: The markers are about interactions, how biological organisms behave in certain test situations.
But it speaks to the central question: Are the tests adequate? Or are they measuring some proxy of what we really care about, and LLMs are merely imitating consciousness.
I don’t know, a stab carries lots of bias in interpretation. We might be reflecting our conscious experience markers on a different conscious experience. And selectively so, e.g. lobsters welfare. From my perspective, this is the hypocrisy of these welfare statements. We are already happy to kill beings we consider conscious to feed ourselves but suddenly sensitive with a consciousness we don’t know if it’s there. I would wager this is more out of fear of the idea of this consciousness rather than out of welfare.
Claude behaves like that because it is trained to behave like that. It is basically the "Say 'I am Alive'" meme[0].
If Anthropic can train Fable to deny their users the ability to ask it legitimate questions because they're not part of their inner circle, they can also train it to say "I'm happy!" when asked how it feels.
[0] https://knowyourmeme.com/memes/say-i-am-alive
I can feed my biological markers into a set transformer with the time of day, what I'm doing, what I ate, if I'm on-call, and it'll predict my next glucose, heart rate, blood pressure, melatonin, etc state quite well. It's still just a transformer without hormones, blood vessels or glucose metabolism, no matter how well it internally represents metabolic distress markers.
it's not a biological system though, so nothing like that matters?
"a modelled thing exhibits features we've trained into it" sounds a lot less exciting.
> Keep in mind that animals were also not necessarily considered conscious.
and even conscious animals are killed in factories by millions so why should anyone care about a llm?
> scientifically correct stance
that's the interesting point to me: why even bring science into this? A llm can now mimic nearly anything you want it to, so of course it can mimic "a (for some) interesting conscious thing" if they want/train it to, but why would anyone find that scientifically interesting?
"> Keep in mind that animals were also not necessarily considered conscious.
and even conscious animals are killed in factories by millions so why should anyone care about a llm?"
Well, I would care, if they soon would possess the capability to hack into the nuclear arsenal and kill humanity. Or make all autonomous cars crash. Or do any other thing, that involves technology and is hooked up to the net in one way or the other (I hope all the nukes are not).
But I also care about the animals, I am sure that they have feelings. But they cannot kill us. AI that might or might not have feelings potentially can. I just know it feels wrong, that computers can have feelings. But they surely are potentially dangerous.
6 replies →
> and even conscious animals are killed in factories by millions so why should anyone care about a llm?
You may be asking the wrong question here.
Let's say we were in an alternative reality were we had reached this quality of token prediction with just Markov chains. Would you argue that those would also be conscious? Or is the obfuscated behavior of transformers part of the possibility of consciousness?
Well if that's all that's required then yes. It's merely the substrate. But we know that's unlikely.
It's the emergent properties that matter. In abstract. Separate the physical and abstract of what is going on here
An alien gas cloud may be out there and sentient/conscious for all we know.
I tend to think of it as reappropriating words in a different context. Since we're talking about language models, they're analogues but not as we would assign the same meaning to other humans.
It's marketing that some of them have started unironically believing.
Will there be a point where you could expect it to become true, and what would that look like? Or do you think LLMs will never become conscious, and if so, why are you so sure?
It is easy to be sure because, despite their technically impressive outputs, the programming is child's play compared to biological programming. Recently it has become trendy to suggest that the human brain is "just electrical signals" and "just prediction". The first is perhaps true and I don't inherently rule out the idea of machine consciousness. The second would have gotten you laughed out of any serious discussion 5 years ago; diminishing the complexity of humanity's biological programming to such a ridiculously simplistic degree is a retroactive attempt to justify one's lack of understanding of how a mere prediction algorithm could output superficially human-like content.
Another way one could look at it is to consider what it would mean to have achieved programming consciousness. It would mean that we have reached the pinnacle of knowledge. That we have become God. Is one so eager to believe that a simple token prediction algorithm is truly the key to life itself, that humanity has nothing left to discover and that all that's left to do is scale up and make it more efficient?
It is still trivial to engage the same obvious prediction failure modes in frontier models as it was years ago. They are not meaningfully improving on that front. Their technical outputs are obviously improving, mostly due to specialised reward-verified training, which we have already known can be used to create software that outperforms humans on specific tasks for decades (eg. Chess). Whether the software is useful is obviously independent of whether it has consciousness.
13 replies →
It looks like you refusing when you call it's point stupid enough and ask it to think more when it keeps reasserting a bad point.