Integrative Turing-like tests for Language and Vision
2022 verified 4.1.0 Mengmi Zhang, Elisa Pavarino, Xiao Liu, Giorgia Dellaferrera, Ankur Sikarwar, Caishun Chen, Marcelo Armendariz, Noga Mudrik, Prachi Agrawal, Spandan Madan, Mranmay Shetty, Andrei Barbu, Haochen Yang, Tanishq Kumar, Shui’Er Han, Aman Raj Singh, Meghna Sadwani, Stella Dellaferrera, Michele Pizzochero, Brandon Tang, Yew Soon Ong, Hanspeter Pfister, Gabriel Kreiman · arXiv preprint
We systematically benchmark AI’s ability to imitate humans in three language tasks (image captioning, word association, conversation) and three vision tasks (color estimation, object detection, attention prediction), collecting data from 636 humans and 37 AI agents. Next, we conducted 72,191 Turing-like tests with 1,916 human judges and 10 AI judges.
Format: Displaced Modality: Multimodal Threshold: Chance (50%)
Run against a system Extensive — the strongest empirical entry in the corpus alongside TuringBench, Rahimov et al., and the Generalized Turing Test. For the Conversation task: judges correctly called genuine human exchanges “human” 66% of the time, and correctly called AI-generated exchanges “AI” only 47% of the time; overall imitation-detectability was 0.57 for conversation (0.57 image captioning, 0.53 word association) — close to chance. In AI-AI conversation pairs, Blenderbot exchanges were judged human 67% of the time, more often than genuine human-human exchanges (64%). A simple SVM judge trained on single sentences matched or beat human judges at catching machine text despite far less context. Separately, the authors ran a live, interactive, 3-party version (one judge, one human agent, one AI agent, judge asking questions in real time) and found human judges reached up to 100% accuracy in an early pilot, degrading toward chance only as the number of allowed exchanges shrank — a striking contrast (offline/displaced judging near chance vs. live/interrogative judging near-perfect) directly relevant to this corpus’s thesis that format changes what is measured.
How this was mapped onto the axes, and where the paper is silent
A genuine boundary case, and the call is close. The paper’s headline contribution is a large-scale, offline, forced-choice test: crowd-collected human and AI responses to 3 language tasks and 3 vision tasks are shown to judges who pick which of two responses is human. Only one of the six tasks (conversation) is conversational in Turing’s sense; color estimation and object detection are psychophysical judgments with no dialogue. Mapped as non-fork because the underlying evidentiary method is identical across all six tasks and identical to Turing’s own move: behavioral output, forced-choice human-vs-machine judgment, nothing mechanistic — the paper generalizes indistinguishability judgment across modalities rather than replacing it with a different construct (contrast with the NeuroAI Turing Test in this corpus, which swaps behavior for internal representations). format=displaced reflects the primary, large-N test: judges read/view pre-collected transcripts and vision stimuli rather than interrogating live. The paper also ran a much smaller supplementary live 3-party experiment (one judge simultaneously conversing with one human and one AI agent, GPT-3.5-Turbo) that gave very different results (see empiricalResults) — not the paper’s primary config, but a striking finding about how format changes what is measured. duration is defaulted: the paper caps conversations by exchange count (up to 24, extended to 48), not minutes. threshold=chance is an editorial call: the authors never declare a formal numeric pass bar but repeatedly interpret their “imitation detectability” scores (0.53–0.57) against an implicit chance/0.5 reference. rigor is defaulted to informal despite substantial disclosure (public code/data, cross-platform replication, corrected permutation tests), since no pre-registration is mentioned. Version note: v1 (2022-11-23, the date recorded above) already contained the core methodology but with smaller numbers (769 human agents, 24 AI agents, 21,570 trials); the abstract quoted above reflects v3 (Sep 2025) figures. The live-interrogation sub-experiment using GPT-3.5-Turbo postdates v1, since that model did not exist until March 2023.
Can Machines Imitate Humans? Integrative Turing-like tests for Language and Vision Demonstrate a Narrowing Gap · doi 10.48550/arXiv.2211.13087
The Turing Test Is More Relevant Than Ever
2025 verified 1.2.0 Avraham Rahimov, Orel Zamler, Amos Azaria · arXiv preprint
This study argues against discarding the Turing Test, proposing instead using more refined versions of it, for example, by interacting simultaneously with both an AI and human candidate to determine who is who, allowing a longer interaction duration, access to the Internet and other AIs, using experienced people as evaluators, etc.
Interrogator: Motivated amateur Threshold: Chance (50%)
Run against a system Real and central to the paper. Section 4, Tables 1–3: four between-subjects MTurk conditions crossing {Simple, Enhanced} × {no prompt engineering, with prompt engineering}, using Llama 3.2 1B as the AI throughout. Results: Simple/No-Prompt, n=41, 68.29% correct identification; Enhanced/No-Prompt, n=29, 93.10% correct (chi-sq(1)=4.97, p=.026); Simple/With-Prompt, n=41, only 43.90% correct — the prompt-engineered LLM was misidentified as human more than half the time, read by the authors as the AI “passing” this weaker version; Enhanced/With-Prompt, n=31, 70.97% correct (chi-sq(1)=6.21, p=.013) — the same LLM clearly failed the harder, three-party version. This directly supports the abstract’s central claim with real numbers: an LLM can pass a weak version of the test but fails a more robust one. No independent replication of this specific Simple-vs-Enhanced protocol found.
How this was mapped onto the axes, and where the paper is silent
Important correction to my own working assumptions: the abstract’s laundry list (longer duration, Internet access, expert evaluators) describes the paper’s aspirational “Ultimate Turing Test / Turing Test 2.0” discussed in Section 6, which is not what was actually run. What the paper actually implements and tests (Section 3) is narrower: a “Simple Turing Test” (two-party, single chat window, 2-minute limit) versus an “Enhanced Turing Test” (three-party: a tester converses simultaneously with a human responder and an AI in two separate windows, exactly Turing’s 1950 structure, 5-minute limit). Internet access and domain-expert evaluators are proposed as future work, never implemented; extended duration also never appears in the tested design — the Enhanced Test’s 5 minutes matches, rather than exceeds, the 1950 baseline. Config maps to the Enhanced Test, the paper’s actually-validated, recommended design and central empirical claim. format=three-party because the Enhanced Test’s dual-chat, simultaneous-comparison structure matches dimensions.json’s three-party definition directly. interrogator=motivated rather than naive: testers are Mechanical Turk workers screened for quality and paid a bonus specifically for correctly identifying the human/AI — financially incentivized to win, though not domain experts; a reader could push this toward naive instead. threshold=chance: the paper interprets its results relative to a 50% cutoff throughout (e.g., treating 43.9% correct-identification as the AI having “passed”). rigor=informal: real statistics (chi-squared tests) and full model/prompt disclosure, but no pre-registration.
The Turing Test Is More Relevant Than Ever · doi 10.48550/arXiv.2505.02558
Dual Turing Test
2025 verified 2.2.0 Alberto Messina · arXiv
In this short note, we propose a unified framework that bridges three areas: (1) a flipped perspective on the Turing Test, the “dual Turing test”, in which a human judge’s goal is to identify an AI rather than reward a machine for deception; (2) a formal adversarial classification game with explicit quality constraints and worst-case guarantees; and (3) a reinforcement learning (RL) alignment pipeline that uses an undetectability detector and a set of quality related components in its reward model.
Format: Displaced Interrogator: Motivated amateur Threshold: Chance (50%)
Never run None. The paper self-describes as a “short note” and is purely a formalization. No dataset, no participants, no trained detector, no reported accuracy numbers; Section 8 (“Proposed Immediate Actions”) explicitly lists building a pilot benchmark as future work — it does not yet exist. The RL alignment pipeline (Sections 4–5) is likewise presented as a proposed training loop with no runs reported. No independent implementation found.
How this was mapped onto the axes, and where the paper is silent
The judge stays human throughout the paper’s own contribution (Section 3): a fixed judge is presented, per round, with an unlabeled pair of replies — one human, one AI — to a sampled prompt, and must say which is which. The “Inverted Turing Test” the abstract mentions is explicitly cited as prior art (Watt’s classic variant where a machine judges), not this paper’s own mechanism — so inverted (machine judges which interlocutor is human) does not apply; this is still a human-judge game with the payoff flipped from being deceived to catching the deception. format=displaced is a judgment call: prompts are drawn from a fixed prompt space rather than authored/adapted live by the judge, and the judge only classifies a pre-generated reply pair rather than conducting an interrogation — closer to reading a transcript than steering one; a reader could argue for three-party instead since both outputs are presented side by side for one comparative verdict. threshold is the hardest fit: the paper’s own pass criterion is a minimax detection accuracy for the judge (≥ 0.70 in one worked example), the mirror image of this schema’s deception-rate framing; its one explicit statistical anchor is a binomial test against a 50% chance baseline, which is what chance is mapped from here, but no option in this schema was built for a flipped pass criterion. interrogator=motivated reflects that the entire premise is a judge deliberately trying to catch the AI, not Turing’s naive general-public interrogator, though no domain-expertise or tool-use is specified. Sections 4–5 (an RL alignment pipeline using the detector as a reward-model critic) are a separate training-procedure proposal layered on top of the dual test, not part of the test itself — noted as a design element the schema doesn’t capture, but not enough to push this into fork territory since Section 3’s core mechanism is still a human-vs-AI discrimination game.
Dual Turing Test: A Framework for Detecting and Mitigating Undetectable AI · doi 10.48550/arXiv.2507.15907
Energy Efficient Imitation Game
2025 verified 2.0.0 Adam Winchell · arXiv
This work expands upon the original imitation game by accounting for an additional factor: the energy spent answering the questions. By adding the constraint of energy, the new test forces us to evaluate intelligence through the lens of efficiency, connecting the abstract problem of thinking to the concrete reality of finite resources.
Resource constraint: Energy budget
Never run None — a philosophical essay with no experimental section, run against no real system. The central instrument, the “psychoergometer,” is explicitly described as nonexistent: “psychoergometers do not exist; and if they did, it would, for obvious reasons, be impractical to use them at scale.” There is no Pareto frontier, no measured joules-per-answer data, and no deception-rate-vs-energy tradeoff plotted anywhere in the paper — verified by full-text search: Pareto and joule each occur zero times in the document. The paper is an argumentative essay (motivating example, formal game description in prose, a “Contrary Views” section, closing thoughts on a hierarchy of intelligence and the Halting Problem), not an empirical study. No independent implementation found.
How this was mapped onto the axes, and where the paper is silent
A genuine same-game extension, not a fork: the paper explicitly keeps and re-derives Turing’s three-party structure (“It is played with three players: the liar, the truthteller, and the interrogator… What will happen when a machine takes the part of the liar in this game?”) and states, “We can actually embed the original imitation game within this new game by asking ‘Can machines think?’” — so format=three-party and constraint=energy are both grounded directly in the text, the latter being the paper’s whole contribution. Everything else is left unspecified by the paper and defaulted: no channel restriction, no session-persistence discussion, no time bound for the interrogation itself, no interrogator-sophistication spec, and no numeric pass criterion is ever given. evidence=behavioral rather than mechanistic: the energy measurement (via a hypothetical “psychoergometer”) is an external property of the answering process, not an inspection of internal representations.
Can machines think efficiently? · doi 10.48550/arXiv.2510.26954
A Rigorous Turing Test
2025 verified 1.4.1 Sharon Temtsin, Diane Proudfoot, David Kaber, Christoph Bartneck · arXiv
Several studies claim that large language models have passed the Turing Test and hence can “think”, yet none follow Turing’s original instructions precisely. Passing the test holds significance as evidence that a machine demonstrates human-like intelligence, and as a marker for artificial-general intelligence in commercial and legal domains.
Duration: Unbounded Threshold: Chance (50%) Methodological rigor: Pre-registered
Run against a system Yes — the strongest, most rigorously reported empirical run in this corpus. Pre-registered, IRB-approved. Recruited 210 participants across 42 CIHG and 42 man-imitates-woman benchmark game (MIWG) trials; after exclusions, 185 participants remained across 37 trials of each game type. Machine witness: GPT-4-Turbo via the OpenAI API (temperature 0.75, max tokens 1000, frequency penalty 0.0). Headline finding: in the CIHG, only 1 of 37 interrogators misidentified GPT-4-Turbo as human — 36/37 (97%) correctly identified it as non-human, significantly above the 50% chance level (p<0.05, 95% CI 88–100%). Mean CIHG duration ~14 minutes. The MIWG benchmark produced near-chance identification accuracy (~43%, 95% CI 28–59%, not significantly different from 50%) — interrogators could not reliably tell a man imitating a woman from an actual woman via text either. The paper’s conclusion: “GPT-4-Turbo did not pass the test,” directly contesting prior “LLM passed the Turing test” claims on methodological grounds (those studies used time limits Turing never specified). Note: the published Results section contains one internally inconsistent sentence that appears to transpose the 43%/97% figures between the two games, contradicting both the abstract and the adjacent sentence; this was resolved by checking the paper’s own reported odds ratio (42.86), which only reconciles with CIHG=97%/MIWG=43% (the abstract’s reading), confirming that one inline sentence carries a labeling error.
How this was mapped onto the axes, and where the paper is silent
Not a fork, and arguably the opposite of one: a strict, literal replication of Turing’s 1950 three-party game, run more carefully than prior claimed “passes,” not a redefinition of what is measured. format=three-party and evidence=behavioral are explicit and central (“We conducted Turing’s three-player imitation game… following the guidelines identified by Turing”). modality=text and memory=single are explicit in the described setup (single-sitting, message-based game via the OpenAI API; each participant took part in only one trial in a single role). duration=unbounded is explicit and is the paper’s central methodological point: the computer-imitates-human game (CIHG) was run “without duration constraints,” and the authors argue this is precisely why they get a different (negative) result than time-boxed prior studies. interrogator=naive is explicit, not defaulted: the paper quotes Turing’s own spec that the interrogator “should not be expert” and confirms recruited participants were not required to have AI expertise. constraint=none is defaulted: the paper never imposes a compute or energy budget. threshold=chance: the paper never adopts Turing’s 30% figure as its criterion, and its own significance test of interrogator accuracy is against a 50% chance level. Its preferred benchmark, interrogator accuracy in the man-imitates-woman game, has no option on this scale, and the paper itself declines to recommend chance, 70% or that benchmark over the others. rigor=preregistered: preregistered on AsPredicted (#148990) and IRB-approved (Univ. of Canterbury HREC 2023/98/LR-PS), with the model (GPT-4-Turbo), API parameters and full system prompt disclosed; it reports no compute figure, which the full option requires. Title note: the current arXiv title is “A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence” (lowercase “a”), which differs from the working title “The Imitation Game According To Turing” found in earlier secondary coverage — this is the live, current title, recorded verbatim.
A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence · doi 10.48550/arXiv.2501.17629
The Generalized Turing Test (GTT)
2026 verified 2.1.0 Daniel Mitropolsky, Susan S. Hong, Riccardo Neumarker, Emanuele Rimoldi, Tomaso Poggio · arXiv
We introduce the Generalized Turing Test (GTT), a formal framework for comparing the capabilities of arbitrary agents via indistinguishability.
Format: Two-party Threshold: Chance (50%)
Run against a system Yes — a real, substantial empirical section, self-described by the authors as exploratory rather than a rigorous measurement campaign. Nine LLMs tested pairwise: Claude Opus 4.6, Claude Sonnet 4.6, GPT-5.4, Gemini 3.1 Pro Preview, DeepSeek-V3.2, Mistral Large 2512, Ministral 8B 2512, Qwen3 32B, and Grok 4.20. For each of two full-matrix protocols (plain GTT, and GTTQ which adds a querying phase), every one of the 81 ordered actor-distinguisher pairs was run 10 times, giving 810 trial records per protocol — at least 1,620 trials across the two main protocols alone, plus additional fixed-distinguisher experiments, consistent with the abstract’s claim of “thousands of trials.” Headline result (Table 1): ranked by mean aggregate Turing Score T, Gemini 3.1 Pro tops the table (T=0.784, F=0.750, D=0.819), followed by Opus 4.6 (T=0.734), GPT-5.4 (T=0.722, but lopsided: F=0.912, the best fooling/actor score in the table, against a comparatively weak D=0.531 as distinguisher), Sonnet 4.6 (T=0.678), DeepSeek V3.2 (T=0.603), Grok 4.20 (T=0.569), Mistral Large (T=0.478), Qwen3 32B (T=0.450), Ministral 8B lowest (T=0.428). The authors report this ordering is “broadly consistent with familiar model leaderboards.” A controlled-turn ablation found a strong dose-response effect for a strong distinguisher: “against Gemini imitating Claude Opus, the distinguisher’s success rises from 40% at one turn to 90% for ≥3 turns.” No human baseline or significance testing is reported for the main pairwise matrix; treat per-model numbers as descriptive, as the authors themselves caveat.
How this was mapped onto the axes, and where the paper is silent
This is the proposal in the corpus that strains the nine-axis schema hardest, because it removes both fixed roles Turing assumed. Formally: A ≥ B iff B, acting as “distinguisher,” cannot reliably tell an interaction with a fresh instance of B from an interaction with A instructed to imitate B. Neither A nor B need be human; the paper’s own worked example is Gemini 3.1 Pro imitating, and being distinguished by, Claude Opus 4.6. The word “interrogator” does not appear anywhere in the paper (checked by full-text search) — “distinguisher” replaces it deliberately. FORMAT: each trial is one distinguisher conversing with one unknown interlocutor and rendering a binary verdict, with no simultaneous side-by-side pair and no separate human judge — structurally closest to two-party (a witness judged in isolation), just with a non-human witness and non-human judge. MODALITY: interactions are explicit multi-turn LLM conversations, so text is not a default. MEMORY: each trial is fresh and bounded, matching single. EVIDENCE: purely behavioral — the distinguisher only sees transcripts, never weights or activations. DURATION genuinely does not map: trials are paced in turns (main GTT capped at 40 distinguisher turns, GTTQ query phase at 20), not minutes; left at the 5min default for lack of a better option, but a 40-turn LLM exchange is almost certainly far more content than a human 5-minute teleprinter session — the schema needs a turn-count axis to represent this proposal honestly. INTERROGATOR is the single biggest mismatch: the distinguisher in every reported experiment is a model instance, never a human. None of naive/motivated/literate/expert describes “the judge is itself an AI system”; defaulted to naive only because there is no non-human option. THRESHOLD mapped to chance: the pass criterion is Pr[B succeeds] ≤ 1/2 + ε (ε = 0.005 in the main empirical figure) — the distinguisher must do no better than a coin flip within a small tolerance. A stronger “statistical indistinguishability” definition is also given but explicitly described as impractical to test and not what the empirical section operationalizes. RIGOR: informal, matching the authors’ own characterization: “this study is intended as a first empirical instantiation of the framework rather than a high-precision measurement campaign,” no preregistration, no significance testing on the main pairwise matrix.
The Generalized Turing Test: A Foundation for Comparing Intelligence · doi 10.48550/arXiv.2605.10851