Comment by blurbleblurble

13 hours ago

"we demonstrate that LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist"

It's almost as though bias-making machinery is embedded in the texts these things are trained on.

It's wild to see quantitative researchers catching even just a glimpse of what culture/media/literary theorists have been swimming in for decades.

"under realistic conditions, contemporary LLMs ... consistently favor otherwise identical resumes with female or stereotypically black names over those with male or stereotypically white names, even when explicitly prompted not to show any race or sex preferences"

https://arctotherium.substack.com/p/llm-fairness-in-realisti...

  • Don't hold your breath on "culture/media/literary theorists" mentioning that. Nor the fact that "the odds of success were identical for every group at every job" is a completely unrealistic assumption.

No, the bias-making machinery is embedded in the machinery, part of the purpose of which is to do a rough kind of statistical analysis via "attention". If for example "Tufa" keeps appearing (n=small, but more than for the other fake tribes) next to terms indicating skill at some task, of course that will be noticed. It will last for as long as that information is in the context window (weights for the current model don't get updated as a result of conversation; that's just not how they work). And of course that can happen from random chance, and of course the LLM has no way to externally verify the extent to which randomness is in play (or the ground-truth probabilities).

The paper makes clear that they used pre-trained, frontier models — in other words, they did not train models on fake data about the fake tribes that would ascribe fake stereotypes to them. There is nothing to suggest that the training data somehow accidentally encoded biases related to fake tribes that the creators of the training data (i.e. ordinary people going about their ordinary Internet lives) somehow accidentally expressed.

There is also nothing to suggest that reading the entire Internet would somehow predispose the reader towards the general idea of being "biased", in the sense that you would have to have in mind to see an actual problem here. But really, the kind of "bias" we're talking about here is really pattern-matching on the available data, which is a big part of what leads people to apply the term "intelligence" to the models. See also the way that people try to make "culturally neutral" IQ tests specifically by having them focus on the ability to infer patterns (e.g. https://en.wikipedia.org/wiki/Raven's_Progressive_Matrices ).

  • I think you're slightly misunderstanding my point, which is that a huge spectrum of associative pattern seeking logics are embedded in language and that the LLM learns them and operates them, approximately.

    "There is nothing to suggest that the training data somehow accidentally encoded biases related to fake tribes that the creators of the training data (i.e. ordinary people going about their ordinary Internet lives) somehow accidentally expressed."

    This is exactly not what I'm suggesting.

    • I think it's less obtuse to make the point that there are social-prejudicial structures encoded in the latent space. "associative pattern seeking logics" is a bit incoherent, and distributional semantics doesn't live at a level accessible to cultural analysis and theoretics, IE film critique.

      If you walked into a film theory class and posited that you could derive every single encoded interpretation of a film by memorizing the positioning of the actors and objects frame-by-frame, plus the audio track in another language that you do not speak, for every single piece of video ever made, you'd probably be asked to leave. Not that I'm advocating for the position of critique here, I don't think the anti-distributional semantics crowd is ever going to recover from their humiliation that's been accelerating over the last 8 years. It's just that from the position of critique it requires a coherent narrative that human brains are capable of ingesting (IE not maximal information overload).

      If you really want to go the lower level route, I think Francois Laruelle's non-philosophie touches on what you might be thinking of in a much more robust way, shining a light on the unexamined consequences of decision and dialectics of-themselves. If you can stomach the writing of continentals, that is.

      1 reply →

  • It could be that a model prefers the tribe mentioned in closest proximity to the word candidate most of the time. It could be that it prefers the one that's third in a series. It could be that it prefers the one with even numbers of letters.

    The model is biased. That's it's entire function, to bias certain tokens over other tokens based on a bunch of vectors and context. There's no telling what is influencing that bias.

    The models will be statistically more likely to choose one of the options for completely unknowable reasons.

I think that quantitative researchers have known this for a while, too.

My perennial experience as a machine learning practitioner working in industry is that the ML and statistics folks raise concerns about the models learning social biases that could case real harms, the business folks make sure that this is a career-limiting move, and so the quantitative folks learn not to rock the boat.

> "we demonstrate that LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist"

For a while (It's getting better with Astra, but still there), a lot of these models would "accuse" you of wishing that magic existed or something, and constantly drawing distinctions to try and "prove" something that nobody ever said.

I think that holding and generating distinctions, when it comes to problem solving, is a very powerful tool. If nothing else, it's a way to force yourself to be adversarial. Conflation is a "damning" operation, while distinctions will at most blow up your search complexity (which, we know from computer science, isn't free, but still).

But it's not a way to build a model, a theory, a society. It's like permanently being the "uhm, actually" redditor.

This is not only a fairness problem. It is an agent-memory problem: a system can mistake its own early choices for evidence.

The whole abstract is full of falsehoods and unsubstantiated assumptions, dare I say unjustified biases.

There have been a few papers recently suggesting that ChatGPT responds differently to different demographics. Specifically, depending on your gender, education level, socioeconomic status, race, and other characteristics, or how it reads those, it might give less accurate responses to the same prompts. These unfavorable outcomes are generally unfavorable in the ways that one would expect of course

https://www.sciencedirect.com/science/article/pii/S187705092...