Comment by adsharma
4 hours ago
The real threat is that we uncritically adopt language such as alignment.
Implicit in this is the idea that AI is a inscrutable matrix and going to remain that way and we'll need expert interpreters to make sense of it.
We need to insist on building tech that's explainable by design.
Alignment just means "this machine operates in ways that align with the intent of its users". It doesn't imply anything about the inscrutability of the machine in question. A gun with a misaligned scope would likewise fail to operate in accord with its user's intent, and likewise with potentially deadly consequences.
The gun comes with a manual on how to use it safely. I'm sure it has some complexities, but at the high level:
> A gun is a metal tube that uses a tiny, controlled explosion to shoot a small piece of metal (called a bullet) forward at very high speed.
If the gun doesn't work as intended, you can take it to a shop and someone can fix it so it works as designed.
All I'm saying is AI should be designed the same way. Treat AI as normal tech like any other and use similar language.
Learning ML, there was a high emphasis on the error part of things as most of the course was on minimizing errors. After ChatGPT, there is a weird anthropomorphization going on, where it's all about hallucinations, alignment and what not.
We have something that is statistical in nature so there should never been any expectation of error-free results/actions. The value has always been about discerning trends or the cost of errors being way lower than any good result.
1 reply →
That last line needs a lot of workshopping. A guillotine with instructions on the bottom of the blade conforms to your request.
If there is guillotine in the weights of the model, it needs to be properly labeled so you can look it up by name using a database index (or a graph-vector index).
It helps both the bad guys and good guys. Like responsible disclosure in cyber security, we need to have a conversation around it.
> We need to insist on building tech that's explainable by design.
You realize this means insisting on terrible tech that humans can understand right? It essentially caps human progress at some point about 4 years ago.
If you are old and happy with the way things are this might sound like a good idea. It does not to me.
Oh you want tech that helps discover new science instead of parroting existing wisdom?
There is little evidence that the RSI we are discussing is capable of inventing the theory of relativity (or the more advanced equivalent). All we have seen is pattern matching in a much larger space than humans can, with some human provided verification tech.
I would argue that human-AI collaboration with explainable tech has a better chance. Continuous learning can be done in a way that doesn't violate IP or privacy.
You are assuming that humans are capable of understanding everything. They are not.
Human comprehension sets a ceiling on progress.
It reminds me of those schools that can only teach as quickly as the dumbest kid in the room can follow. We don't want that for our entire species.
1 reply →