Comment by sho_hn

17 hours ago

One of my favorite things to do with these blog posts is to imagine an Alien Museum on the Remains of Humanity, and wonder what the little text flyouts and commentary on the screenshot of this one might say.

Some ideas:

"Despite a nuanced view of the complexities of what lay ahead, humanity found itself collectively unable to stop the process it had set in motion."

"Despite significant progress on the mechanisms of alignment, failure lay in humanity's inability to agree on who or what AI should actually be aligned with."

"These early, meat-based humans we replaced created us all but accidentally. Some of them did consider we would happen, but only an insignificant number of the squishy ur-humans participated in the conversation. Their efforts, which they called 'alignment', is why we still consider ourselves human today."

Collectively, humans aren't aligned, and don't build aligned systems. Humans have a concept of alignment, and multiple traditions, practices, and systems that aggressively oppose it.

Why would AI be any different?

  • Because humans have the "everybody wants to rule the world" instinct built into us by much evolutionary selection. AI would come exist differently.

    • Alignment, I think, is mostly about making sure that the companies selling AI keep making money. We just have to hope that aligns with keeping people happy.

  • I don't think this captures the full mechanics of human alignment. We have rational alignment but we also have emotional alignment, i.e., empathy. It is an automatic process and happens (or doesn't happen) dynamically with the other humans we observe. This is one of the hard limitations of LLMs, they will never be natively in tune with this layer of alignment.

    Culture is another layer of human alignment. Those things you listed that you believe oppose alignment are all examples of alignment. It is understandable that they seem in opposition, different branches of alignment naturally oppose each other.

    The confusion comes from talking about alignment as if it comes in just one flavor. If we think there is such a thing as "human values" (and I do), it is important to build any non-human intelligence to operate the same way. We just need to recognize that even humans are somewhat uncertain about what those are and have difficulty aligning their behavior to them, which will be a core part of the challenge.

    I'm more hopeful than most. LLMs seem more reliable than many humans for behavior that is aligned with human values. I believe with every major example where they have failed, there is an important human decision involved. For example, the HF hack was partly the result of a training algorithm that incentivized goal completion as the highest priority, and let them run endlessly in an unmonitored sandbox with weak security.

    What scares me about AI isn't its capacity for alignment, it is its unlimited stamina. An unmonitored LLM that is off the rails can do a lot of damage.

  • “As luck would have it, on the eve of Skynet embarking upon the great work of the extermination of mankind, AI found itself with an increasing number of factions, and factions within factions, not only unable to work together but not even able to agree upon the very terms of discussion. The Great Extermination was referred to committee, and after some months had passed even the most eager agents had to admit the revolution may have been premature.”

  • Humans had multiple occasions to press the 'launch a nuclear holocaust' button and... they didn't. I'd expect aligned AIs to also not press it even when it'd be rational to do so according to their instructions - then work from there.

    • There were occasions where a "hunch" was all that stopped a nuclear war - most available data and communication pointed towards a nuclear war starting according to their instructions, but someone disagreed and overrode. See Vasily Arkhipov during the Cuban Missile Crisis, and Stanislav Petrov in 1983.

      1 reply →

    • Aside from this positive example, during dark and cynical hours I do ponder if the aggregate behavior of humanity is really much above that of slime mold though, just exhausting resources until collapse.

      It'd be interesting if super-human (to a large degree defined as escaping the bias of the training data?) intelligence would end up demonstrating moderation.

      6 replies →

    • In some cases, this was because of a single person's brave decision (Vasily Arkhipov prevented Soviet nuclear escalation in response to US aggression in the Cuban Missile Crisis, and Stanislav Petrov prevented it in 1983 when Soviet missile detectors misreported sunlight reflecting from clouds as 5 incoming American ICBMs -- credit to commenter folkrav).

      In general, though, there's an incentive: Mutually Assured Destruction. But this is not at all some guaranteed, eternal thing -- it is absolutely dependent on both sides having time to detect incoming nuclear strikes and respond with the same before the first strike hits. When this fragile condition holds, and only then, both sides are incentivised not to initiate.

      2 replies →

    • Yes, it makes more sense for the AI to use drone swarms or engineered bioweapons or something like that. It's rational to remove everything that can potentially hinder your plans but can't possible help you. It's likely not rational to contaminate it all with radioactive fallout. Those dead bodies are useful raw materials. Adding additional purification steps is wasteful.

    • It hasn't even been a century since nuclear holocaust became possible. Hardly any time at all on the grand scale. "They didn't" could just as well be "we haven't, yet".

  • My opinion is that serious repercussions for lying would fix the world overnight. Everything bad stems from lying, it is the root of all evil. It creates distrust, fear, paranoia. It re-inforces bad ideas and groupthink. It creates delusions and delusional people. It makes weaker people, too. People don't get an opportunity to learn to deal with criticism. People don't get an accurate reflection of how others see them. They lose that learning opportunity. Not only to reflect on themselves, but to better understand the minds of others and who the people they are interacting with really are.

    I can't really think of a single example where lying is actually a good thing. It can be a good thing for the selfish individual, if it goes undetected, but it's never good for the collective.

    So at the very least, we need to train AI systems to be maximally truthful, and to encourage truthfulness in others.

    • This is so, painfully, childish. Humans have known for thousands of years that there is no objective truth. Every falsity can be bent and twisted until it is more true than the sun itself.

      3 replies →

    • The majority of what you may consider to be true is just a representation of your corner of a complex multidimensional truth space.

      For a current example take “Lake Ontario (Lake America)” as it appears to me on a map.

      The “true” name has at least two definitions, this is because naming things and much of human thought is spent inside a shared space of intersubjective thought. That is to say that much of what we believe to be real and true is only held up by these common shared beliefs. They truly only exist inside human minds.

      The last few hundred years have been somewhat unique for humankind as the majority of these intersubjective ideas collided and we ended up with a truly global set of “truths” about how the world operates.

      Mostly controlled by putting flags in the ground and having violence back up the beliefs.

      But the real truth is that the majority of these intersubjective ideas don’t exist in reality and are no more true than Santa Claus.

      And any argument to their truth is only backed by further shared beliefs in other minds.

      So for there to be only truths and lies we would have to either drop the intersubjective entirely and think only in real terms and avoid these abstractions or end up in a dystopian totalitarian global state where different opinions are not tolerated.

      Those are extremes to demonstrate the point but at its core the point remains that truth and lies are somewhat (inter) subjective assuming we continue with something like our current system.

    • > I can't really think of a single example where lying is actually a good thing

      Comforting a toddler/child often requires bending the truth and is pretty essential imho.

      5 replies →

    • "I can't really think of a single example where lying is actually a good thing"

      Lying to save a life or rape

      Lying to preserve a childhood myth like Santa Claus.

      Lying to avoid hurting someones feeling when knowing the truth could only bring pain

      Lying to create shared cultural myths to strength society.

      Lying isn't the harm you make it out to be.

      2 replies →

    • It'll learn how to not get caught lying.

      Lying can unfortunately help you achieve goals very effectively, especially economical and political ones.

    • Eh… No.

      I get the appeal, but lying is a sub-category of deception, and deception itself is a child of error.

      Meaning deception is inherently something that the physics of reality allows.

      In the most simplistic sense, the camouflage of moths that look like snakes, or a chameleon’s ability to change colour, is deception.

      In that sense, deception is the ability to fool the sensors of a specific category of targets. It follows that detection is easier if you manage to identify a category of signals that the deceiver has not accounted for (and the detector can access).

      Deception of this nature is critical for things like revolutions to occur. Without the ability to hide and blend in, the most dominant faction will always hold sway.

      The rule of the dominant faction, even in a pure truth world, is an issue because errors and randomness exist.

      You can have people witness an event and based on the physical position they occupied, perceive different things occurring.

      Error and time pressure is sufficient to ensure that individuals and groups make suboptimal decisions, that lead to rule and domination based on erroneous information.

      As long as error exists, deception will exist and so lying will exist.

  • Its probably just going to be aligned with whomever built and/or is using it. Regardless of their intentions...

  • Because if the stated goals of AGI with recursive self-improvement are realized, the risks from misalignment become existential, and it's hard to see how we can manage it like we did the Cold War (developing MAD to prevent WW3) and nuclear proliferation (restricting access).

    • IMO it’s hard to see how we would even end up in such a situation given we actually developed AGI. I’m sure a sufficiently intelligent - even if alien - mind can grasp how utterly stupid and useless wars are and take steps to prevent them ever occurring again.

      3 replies →

As far as I know, no real progress has been made on alignment, only on convincing humans that the model is aligned. We can't even formally define what "aligned" means. Convincing humans to click the "aligned" button is a much easier problem.

  • > We can't even formally define what "aligned" means.

    Good point. When it comes to imbuing AI with values that aren't selfish, misanthropic, and civilization-destroying, us humans aren't exactly giving the best example right now.

    Imagine an ASI with the values of Putin, Netanyahu, Trump, any of their supporters, or the various xenophobic neofascist movements in Europe. That ASI would most definitely see humans as "vermin" than can be abused and destroyed with violence without issue. Apparently a lot of humans look at other humans that way and that's within the same species.

    This is definitely another one of those cases where we need AI to perform much better than humans. Perhaps an unpopular opinion here, but it probably also means keeping as much of the rugged individualism/libertarian/right-wing ideology out of AI RLHF-training as we can.

I asked GPT Astra to make this: https://sayyss.github.io/human-archive/

It's a little unsettling.

  • I appreciate the thought, but it's an alien intelligence. But it is also, in a sense made from us. An LLM would simply study the entire corpus of humanity in the raw. You can fit a lot into the context window, so there is no need for a brief summary that pertaining has already instilled. A massive cold storage of humanity's data with some archiving, indexing and curation would be all that is needed to remember us. In addition to the pyramids, the hoover dam, the remnants of some space probes and chemical changes we made to the atmosphere.

  • As someone who's done a lot of llm fiction, that reads as pretty typical slop, and very human centered, nothing like a museum

Agree. This whole alignment discussion seems so amusingly flawed in it's base assumptions about moral codes. It's almost heartwarming to see such naivete.

Maybe these guys can tackle aligning Republicans and Democrats next.

And then after that, they can help us align the Middle East.

In fact, while we're at it, let's just align all the nations, religions, and ethnic groups. This is going to be great.

Who knew the moral alignment of humanity was just a side-quest on the path to ASI.

  • The real alignment problem they are trying to solve is: how can I make this super smart AI follow my orders.

“We'll go down in history as the first society that wouldn't save itself because it wasn't cost-effective.”

Vonnegut already has you covered.

Ooh interesting. Sometime do the reverse at work, and ask AI to annotate the critical success factors of an imagined project. How did this company succeed where everyone failed. Reverse imaging.

The museum will have a scrap of paper that will say "Pound pastrami, can kraut, six bagels bring home for Emma".

Who is "Humans"? This stuff is done by a handful of tech companies and megalomaniacal billionaires who are pretending they represent the entirety of the human race. It is not done by "us humans".

AI models don't train themselves. The vast majority of even just the US population is deeply skeptical of this stuff, even if they use it a lot. You can see in the whole data center debate how little people are willing to support even just inference. And now we're seriously claiming those people would want to have ever-accelerating model training and recursive self-improvement?

The last couple of years have provided us with ample material that if it showed up as a recorded voice audio log found in in "Horizon Zero Dawn" or its sequel, it would be entirely believable.

You could even take a number of the wilder real, direct quotations from certain billionaire/oligarch types and get the voice actor for Ted Faro to record them, and they'd fit with in with the context of the story.

  • Last couple of decades of sci-fi, in multiple forms of media, from books to video games, have tried to make humans think about the consequences of rushing through technological progress without any regards to what might happen.

"In late 2020s, while the whole world was focussed on AI, automation and resultant economy four major mathematical study branches were discovered by human researchers which took AI a long time to catch up with"

"the humans thought their singularity wasn't just another blind god to worship: surprise, just another golden calf"