Skip to main navigation Skip to search Skip to main content

Large language models pass a standard three-party Turing test

  • University of California at San Diego

Research output: Contribution to journalArticlepeer-review

3 Scopus citations

Abstract

The Turing test has been widely discussed as a test of machine intelligence, but it also provides a measure of how humans distinguish other humans from machines. We evaluated 4 systems (ELIZA, GPT-4o, LLaMa-3.1-405B, and GPT-4.5) in two randomized, controlled, and preregistered Turing tests on independent populations. Participants had 5 min conversations simultaneously with another human participant and one of these systems before judging which conversational partner they thought was human. When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time—not significantly more or less often than the humans it was being compared to. Without these prompts, however, the same models performed significantly worse (38% and 36%), and did not consistently outperform baseline models, ELIZA and GPT-4o (23% and 21%, respectively). A third study replicated these results in 15-min games: two PERSONA-prompted models achieved pass rates of 56% and 59%. The results constitute empirical evidence that artificial systems can pass a standard three-party Turing test. Interrogators’ reasoning focused more on stylistic and socio-emotional aspects of human behavior rather than more traditional notions of intelligence. The results have implications for debates about what kind of intelligence is exhibited by large language models, the social impacts these systems are likely to have, and the aspects of human behavior that people continue to see as unique.

Original languageEnglish
Article numbere2524472123
JournalProceedings of the National Academy of Sciences of the United States of America
Volume123
Issue number21
DOIs
StatePublished - May 26 2026

Keywords

  • AI
  • human-AI interaction
  • large language models
  • social cognition
  • Turing test

Fingerprint

Dive into the research topics of 'Large language models pass a standard three-party Turing test'. Together they form a unique fingerprint.

Cite this