Comment by sublinear

1 day ago

> You'd think that LLMs fulfill this, but I do wonder if a clever human could still discern between them given their particularities.

You don't even have to be that clever. We've all tried using it for work. They fail.

In the first place, the Turing test was always a loosely defined thought experiment. I suppose it's still serving that purpose, but at least half of HN takes it way too seriously. It's absolutely not proof of <engagementbait> endorsed by Alan Turing himself.

I think if you trained an LLM specifically for the purpose of passing the Turing test (instead of being helpful, harmful, and so on), its likely it would pass it.

You would have to train it/finetune it on a couple hundred of 'humans chatting in the context of a Turing test'.

  • I don't think that's true, assuming that the humans are also allowed to strategize and study the problem beforehand. The main weakness of the current models is that they still fail hard at certain kinds of common sense reasoning about real world situations, that no human ever would. Things like the "walk or drive to a car wash that's 2 minutes away" thing from a few months ago (I think the latest models have patched that one in particular, but I'm sure others can be found in a similar vein).

  • That would not really pass the test in spirit, I think. Passing must be something that falls out as a consequence of being a good chatbot, not the other way around.

    However such good conversational agents would still have benefits, and i would be curious how far the approach of building a model that’s nice to talk to can take you, as opposed to the current generation of all-knowledgable helpful assistants.

  • I think you are absolutely right.

    The big problem with the Turing test always was that it doesn't take into account adversarial designs. They had those chatbot contests about a decade ago, where the chatbots would regularly pass the turing test, not because the bots were intelligent, but because they were packed with rhetorical tricks designed to avoid saying anything of substance.

    By any meaningful measure, today's LLM's pass the Turing test. Except we may need to lobotomize them for the deception to work.

  • I'd like to see an honest attempt at that.

    What's more interesting are the types of humans that might fail the Turing test. Maybe to discourage the failure of real humans, there could be consequences outside the context of the test.

    That is, most people accused of being a computer would crack and start pleading their humanity, but some sociopaths might not. I then wonder how many of them overlap with those so invested in abusing the premise of the Turing test.

    It really is a fun thought experiment when you spice it up enough.

Mannerisms aside, they are easy to spot from having inhuman amounts of trivia knowledge

  • Couldn’t you instruct them to not display as much knowledge for a test situation…?

  • Humanity is taking an interesting technological arc.

    Terminator (1984 film) had a scene showing that in a future 2009, humans would use dogs to try to sniff out whether a robot passing for a human is secretly a machine.[1]

    In our real world 2026, there are no humanoid machines that can complete basic generic tasks, like carrying a tray across the stage and holding it for 30 seconds.[2] They move slowly and badly (probably from an LLM like neural network doing very few frames per second of correction and analysis), and are nowhere near lifelike.

    When they don't need a body to pass for a human, such as typing online, they do a bit better.

    We can tell them apart from humans. As you say, they have stylistic quirks. And you mention that they're easy to spot because they come trained with inhuman amounts of trivial knowledge.

    [1] https://www.reddit.com/r/MovieDetails/s/0qUCVPYgjt

    [2] https://www.reddit.com/r/LivestreamFail/s/t6ZV0yhgEe

It won't fool all of the people all of the time. But it succeeds often enough to pass the loose definition. It does many things that we called "AI complete" for decades.

It is also clear that is is also not really aware, either. I think Turing would be tickled that we find ourselves in an intermediate state that he would not have imagined.

I find it more accurate to refer to "A Turing Test" as opposed to "The Turing Test" for that reason. It's just one loosely-defined test; necessary but not sufficient.

I would not call Turing's experiment ill-specified, on the contrary among other factors the paper the game is introduced in earnestly calls for a telepathy-proof room to ensure no side channel leakage. The 1950s was an interesting decade.