Public benchmark
How we measure AI detectors: a 221-document frozen-corpus matrix
This benchmark was measured on a Japanese corpus. It shows how we measure and what we publish — including our own false positives — not a claim about English detection accuracy. Measured August 17, 2026.
Corpus (221 docs, frozen)
90 human-written documents (all pre-2020, provenance-traceable) + 131 AI-generated (Gemini / GPT-4o / Claude, in plain, student-tone, paraphrased, and mixed human+AI difficulty tiers).
Conditions
A score of 70 or above counts as “AI”. Our own chain was measured exactly as it runs in production — the same router, prompts and aggregation, not an approximation.
How to read it
Human columns count false flags (lower is better); AI columns count correct detections (higher is better). “—” means not measured.
| Detector | Human · literary65 docs (pre-2020) | Human · student25 docs | AI · Gemini52 | AI · GPT-4o20 | AI · Claude20 | AI · student tone15 | AI · paraphrased12 (hard) | Human+AI mixed12 (hardest) |
|---|---|---|---|---|---|---|---|---|
| OmniDetect (this site) | 0/65 | 5/25 | 51/52 | 20/20 | 20/20 | 15/15 | 12/12 | 12/12 |
| Pangram | 0/65 | 0/25 | 52/52 | 20/20 | 20/20 | 15/15 | 12/12 | 6/12 |
| User Local每日 100 次/IP 上限,21 篇缺口全在 AI 侧,人写侧完整 | 0/65 | 0/25 | 17/52 | 11/19 | — | 0/15 | 0/12 | 0/12 |
| Sapling | 48/65 | 22/25 | 20/52 | 7/20 | 7/20 | 4/15 | 2/12 | 8/12 |
| ZeroGPT | 0/65 | 0/25 | 2/52 | 0/20 | 0/20 | 0/15 | 0/12 | 0/12 |
| WinstonAPI 实锤不支持日语(LANGUAGE_NOT_SUPPORTED) | — | — | — | — | — | — | — | — |
Our own misses, in the same table
Our chain misjudged 5 of the 25 human student-register documents (omni_score 80–95). Those five are logged as product-improvement targets; we will re-measure and update this page. Publishing the numbers that look bad is what makes the rest of the table worth believing.
