The Detector Disagreement Index
Every AI detector reports a single confident number. Run the same text through three of them and you get three different confident numbers. This page publishes how far apart they actually land, measured on real scans rather than a curated test set. We update it monthly.
Last updated July 28, 2026
Headline figure
Across 287 production scans run between 15 March and 11 June 2026 on the GPTZero + Winston AI + ZeroGPT lineup, the three engines returned different verdict labels on 66.6% of texts. The average gap between the highest and lowest score on a single text was 55.9 points out of 100.
Just under half of all scans — 48.8% — had a spread wider than 60 points. Only 22.3% landed within 10 points of each other.
Cite as: OmniDetect Detector Disagreement Index, July 2026 (n=287, GPTZero + Winston AI + ZeroGPT, 2026-03-15 to 2026-06-11).
How far apart, by pair
Pairwise across the same 287 scans, every combination reached a different verdict on at least two texts in five:
GPTZero vs ZeroGPT — average gap 44.9 points, different verdict on 56.1% of texts, widest gap 100 points.
Winston AI vs ZeroGPT — average gap 36.9 points, different verdict on 48.4%, widest gap 100 points.
GPTZero vs Winston AI — average gap 30.1 points, different verdict on 39.7%, widest gap 100 points.
A 100-point gap means one engine reported near-certain human and another reported near-certain AI, on the same text, in the same scan.
Corroboration from a second lineup
We changed engines during the measurement period, which gives an independent check. The earlier lineup — GPTZero + Originality.ai + Winston AI, 59 scans between 19 February and 12 March 2026 — disagreed on 69.5% of texts with an average spread of 59.2 points.
Two different engine sets, months apart, land within three points of each other. The disagreement is a property of the category, not of one unlucky vendor combination.
We do not pool these numbers into a single figure. Different engines produce different disagreement rates, so a blended average across lineups would not describe anything real.
What we are not publishing yet
Our current lineup — Winston AI + Sapling AI + ZeroGPT, in place since 14 June 2026 — has only 35 scans so far. That is too few to publish as an index figure, so we are holding it until it passes 150.
We are saying this rather than quietly omitting it, because the early reading is higher than the lineups above, not lower. When it reaches the threshold it will be published whichever way it points.
Does this number hurt us?
It is a fair question, so here is the direct answer. We sell multi-engine consensus. Publishing "the engines we run disagree two times in three" can be read as an admission that AI detection barely works.
The honest reply is that the disagreement exists whether or not we measure it. Every one of those 287 texts would have received a single confident score from whichever detector the user happened to open first — and in 66.6% of cases, a different detector would have told them something else. The number is not a property of our product. It is the thing our product exists to reveal.
What it genuinely argues against is treating any single score as a finding. That includes ours when only one engine has run: our own free scan uses one engine, and this data is the reason we tell you not to stop there.
It also sets a limit on what consensus can claim. Agreement between three engines is much stronger evidence than one score, but three engines can share a blind spot — formal academic prose and non-native writing trip several of them at once. Consensus narrows the error; it does not eliminate it, and we would rather say so here than have you discover it on your own document.
Method
Figures come from production scans, not a curated benchmark. Every scan in the sample is a document a real user submitted for their own purposes, which means the distribution reflects what people actually check rather than what a test set was designed to contain.
Verdict disagreement counts a scan as disagreeing when the engines returned more than one distinct verdict label. Spread is the difference between the highest and lowest engine score on that text. Scans are grouped by the exact engine set that ran, and no figure mixes lineups.
We store no document text — only a SHA-256 hash and the detection metadata — so this analysis runs on engine scores and hashes alone. That is also why we can publish it: there is no user content to expose.
Known limits. The sample is self-selected: people scan text they are already unsure about, which likely raises disagreement relative to a random corpus. Sample sizes are in the hundreds, not thousands. And these are the engines we run, not every detector on the market.
