Skip to content
理論シリーズ

How to Read “96.4% Accurate” — the Four Questions That Dismantle Any Claim

What this article gives you: the four questions that expose what any accuracy claim actually means — and worked examples from Turnitin’s own documentation, the RAID benchmark, and a Stanford study.

更新日:August 17, 2026

この記事の証拠ラベル:実測推定解釈立場(実測=検証可能な測定 / 推定=データからの外挿 / 解釈=仕組みの説明 / 立場=私たちの判断)

The four questions

立場

① What was the corpus — academic papers, student essays, marketing copy? ② What was the threshold — at what score did they call something “AI”? ③ Who measured — the vendor on their own test set, or an independent party? ④ Which language was it measured in — an English accuracy figure does not transfer to Japanese or German, as studies keep showing.

A claim that cannot answer all four is weaker than one that can. It is that simple.

Worked example: Turnitin’s “under 1%”

実測

Turnitin states a document-level false positive rate under 1% — but only for documents with more than 20% AI content. The qualifier is part of the number: low-AI documents, which are exactly the hardest to judge, fall outside it.

“FPR under 1%” traveling without its qualifier is a different claim from the one Turnitin actually makes.

Worked example: RAID and the Stanford study

実測

RAID (ACL 2024, arXiv 2405.07940), the largest robustness benchmark, found that most detectors degrade sharply under light paraphrase attacks. Liang et al. (Patterns 2023) found ~61% of non-native English essays misclassified by seven commercial detectors — a figure specific to non-native writing, not a universal FPR.

None of this means detectors are useless. It means every number belongs to its test conditions.

Measured: what we publish

実測

We measure detection channels on a frozen corpus, split by language and register, and publish the full matrix — including our own false positives (5 of 25 student-register documents, currently). Our method and numbers are on the public benchmark page.

Use the four questions anywhere

Open any detector’s marketing page and try to answer the four questions. Claims that cannot answer them are weaker than claims that can — including ours, so check ours too.

関連ページ