Comment by libraryofbabel
1 hour ago
> there’s no way to “crack open” an LLM and see precisely where each skill or tendency lives
Mechanistic Interpretability has entered the chat.
For a classic example, see https://www.anthropic.com/research/tracing-thoughts-language...
The spirit of your point stands, though. This kind of research is interesting to read about, but it's very hard, more like neuroscience or biology than computer science ("LLMs are grown, not made"). You're dealing with a lot of extremely _messy_ complexity, for which organic life is really the only good point of comparison. Most of us here are't really equipped for that kind of work; it's not at all like, say, reverse-engineering a piece of software written by humans. And of course the only people who can do it on frontier models from Anthropic and OpenAI are people within the labs themselves. (But I'm optimistic we'll see more of this work on open weights models...)
No comments yet
Contribute on Hacker News ↗