How the Test Was Run
The team assembled 495 passages from 99 authors, split evenly across blogging, fiction, and scientific writing, all written before ChatGPT launched in November 2022 so the corpus could not have been contaminated by language models. Generation came from Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro; detection from Pangram 3.3.2, GPTZero 2026-05-11-base, and Originality.ai Turbo 3.0.2. The control condition was text generated from ordinary prompts; the experimental condition gave each model five sample passages from an author before asking it to write a new passage in the same style.
Where the Gap Opens
The control condition was no contest: the highest miss rate across the three detectors was 0.7%. On false alarms, Pangram and GPTZero flagged no human passages as AI, while Originality.ai misclassified 19 of the 495. Style imitation changed the picture. Pangram missed 10%, GPTZero 11%, and Originality.ai 18%—roughly 13% on average. Scientific writing was the weakest category, with miss rates of 25%, 24%, and 29% respectively—precisely the genre where these tools see the most real-world use.
How to Read This
Two things are worth separating: this is not a finding that detectors are broken, but that their performance depends heavily on whether the other side is actively evading. Vendor-reported figures typically describe ordinary conditions—GPTZero, for instance, has cited a 99.3% recall figure from a third-party benchmark on its own blog—which is measured differently from Epoch AI's adversarial setup, so the two numbers cannot be used to check each other. The practical impact lands wherever a detector score becomes grounds for punishment. If a school, journal, or content platform treats a single scan as proof of cheating, roughly one in four style-imitated academic submissions will pass, while human authors can still be flagged in error. The safer posture is to treat detector scores as a lead rather than evidence, weighed alongside drafting records, version history, and other signals.
via: Epoch AI's evaluation (data published on its GitHub), reported by The Decoder on July 19, 2026; verified July 21, 2026