Signal Check Research Prototype

AI detection for educators

The arithmetic that makes detection scores unusable as evidence, and what people are doing instead.

Last reviewed 16 August 2026

If you teach, you have a real problem and the tools being sold to solve it do not work. This page is not an argument that the problem is imaginary. It is an argument that detection is the wrong lever, and a summary of what colleagues at other institutions have moved to instead.

Start with the arithmetic

The number that matters is not the accuracy rate. It is the false-positive rate multiplied by your submission volume.

Turnitin advertised a 1% false-positive rate for its AI detection feature. In August 2023 Vanderbilt University disabled the feature and published the calculation behind that decision:

Vanderbilt submitted 75,000 papers to Turnitin in 2022. If this AI detection tool was available then, around 750 student papers could have been incorrectly labeled as having some of it written by AI.

Seven hundred and fifty wrongly flagged papers a year, at the vendor's own advertised rate, at one institution. Every one of those is a student in a meeting, a note on a file, and a relationship with an instructor that does not recover.

Run this calculation for your own institution before adopting any detector. Take your annual submission count, multiply by the vendor's stated false-positive rate, and ask whether you would accept that many wrong accusations as the cost of the tool.

The errors land on specific students

If false positives were randomly distributed, they would still be a problem. They are not random. A Stanford group published in Patterns in 2023 ran 91 TOEFL essays by non-native English writers through seven detectors alongside essays by US eighth-graders:

61.22%
average false-positive rate on essays by non-native English writers
5.19%
average false-positive rate on US eighth-grade essays
19.78%
of TOEFL essays flagged by all seven detectors at once

That is roughly a twelvefold difference by language background. In practice it means your international students, and anyone taught to write in a formal register, absorb most of the error. Any institution running these tools at scale is applying a differential penalty by nationality and language background, whatever the intention.

The same asymmetry shows up in other groups whose writing is deliberately even and conventional: students with certain disabilities, students who use grammar and structure tools as accommodations, and students who were drilled in a rigid essay formula.

Why the vendor accuracy numbers do not transfer

Reported accuracy comes from a benchmark: some corpus of human and machine text the vendor assembled. Your students' work is a different distribution — different genres, prompts, ability levels, first languages, and assignment constraints. Benchmark accuracy sets an upper bound under ideal conditions, not a rate you should expect on your own submissions.

Two further points make the published figures optimistic:

Consider also that OpenAI released a detector for its own models in January 2023 and withdrew it that July, citing low accuracy — it caught 26% of AI text and falsely flagged 9% of human text. If the organisation with full access to the model weights could not make this work, it is worth being skeptical of vendors claiming otherwise.

What to do instead

Assess the process, not just the artefact

Require version history, an outline submitted a week early, an annotated bibliography, or a short reflection on what changed between drafts. This is the single highest-value change: it makes the work visible as it develops, and it is far harder to fabricate convincingly than a finished essay. It also happens to be good pedagogy independent of AI.

Use short oral follow-ups

A five-minute conversation about a submitted piece — why this structure, what you cut, which source changed your mind — separates authorship better than any classifier, and it does not produce false positives. Sampling a subset of submissions rather than all of them keeps this tractable in large classes.

Design assignments that generate specific evidence

Tie work to material a model cannot access: this week's seminar discussion, a specific local dataset, a field visit, an interview the student conducted, an argument made by a named classmate. Generic prompts produce generic essays, which is precisely the condition under which detection looks tempting.

Write an explicit policy and teach it

Most students are not trying to cheat; they are guessing at an unstated line. State what is permitted for brainstorming, outlining, grammar, and translation, and what is not. Ambiguity produces both genuine misconduct and needless anxiety.

If you must use a detector, treat it as triage only

Not as evidence. That means: never open a conversation by presenting the score, never ask a student to prove a negative, never record a finding on detector output alone, and always look for independent corroboration — a change in voice you noticed yourself, an inconsistency with prior work, a citation that does not exist. If the closer look turns up nothing, the score was noise and the matter should end there.

If you are handling an allegation right now

About the tool on this site: the analyzer on the home page has no validated accuracy rate. It should not be used to assess student work, and nothing on this site should be cited in an academic integrity process. We publish it as a demonstration of how these statistics behave, not as an instrument.


Sources

Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7). arXiv:2304.02819
Coley, M. Guidance on AI detection and why we're disabling Turnitin's AI detector. Vanderbilt University Brightspace, 16 August 2023. vanderbilt.edu
OpenAI. New AI classifier for indicating AI-written text (January 2023; discontinued 20 July 2023). openai.com