Humanetext

How AI Content Detectors Actually Work

A plain-language explanation of perplexity, burstiness, and the statistical signals AI detectors use — and why false positives happen so often.

AI detectionBy Humanetext Editorial7 min readUpdated 23 August 2026

The short answer

AI content detectors do not detect AI. They estimate how statistically predictable a piece of text is, using perplexity - how surprised a language model is by each word - and burstiness, how much sentence length varies. Predictable, uniform text is scored as machine-written, which is why clear writing, formal registers, and second-language English are flagged at elevated rates regardless of who wrote them.

AI content detectors get treated like lie detectors — a tool that reveals a hidden truth about how a piece of text was made. In reality, they're statistical classifiers making an educated guess, and understanding how that guess works explains both why detectors catch what they catch, and why they get it wrong so often.

The core idea: predictability

Every AI detector is built around a simple observation: language models generate text by picking the most statistically likely next word, over and over. Human writing is less predictable — we make idiosyncratic word choices, take unexpected turns, and write with more variation than a model trained to produce "safe" output.

Detectors try to quantify that difference using two main signals.

Perplexity

Perplexity measures how "surprised" a language model is by a given sequence of words. Low perplexity means the text closely matches what a model would predict — smooth, expected, statistically average phrasing. High perplexity means the text takes turns a model wouldn't predict, which is more typical of human writing (we use unusual word combinations, personal references, and quirky phrasing that don't follow the statistically "safe" path).

AI-generated text tends to score low on perplexity, because that's literally what the model was optimized to produce: the most probable continuation at every step.

Burstiness

Burstiness measures variation within a piece of text — specifically, how much sentence length and structure fluctuate. Human writing is "bursty": we write a short sentence, then a long one, then a fragment, then a complex compound sentence with three clauses. AI-generated text tends to be more uniform, with sentences clustering around a similar length and structure throughout.

A detector combines low perplexity and low burstiness as its strongest signal that text was likely AI-generated.

Why false positives happen

Here's the problem: perplexity and burstiness are proxies, not proof. They correlate with AI generation, but plenty of things also produce low-perplexity, low-burstiness text that has nothing to do with AI:

  • Non-native English writers often use more predictable, textbook-standard phrasing, which reads as "low perplexity" to a detector even though a person wrote every word.
  • Technical and academic writing is often deliberately uniform and formal by convention, which can trigger the same signals.
  • Heavily edited writing — text that's been proofread and tightened until it's clean and consistent — can lose the natural unevenness that signals "human" to a detector.
  • Short text samples don't give a detector enough signal to be confident, so shorter pieces are especially prone to misclassification in either direction.

This is why every major detector, including the tools built into services like Turnitin, publishes disclaimers that their score is a probability, not a verdict, and shouldn't be used as the sole basis for a decision.

What detectors are actually good at

Despite the false-positive problem, detectors aren't useless. They're reasonably reliable at catching unedited, directly-pasted model output at scale — the kind of text that hasn't been touched by a human editor at all. The uniform rhythm and statistical smoothness of raw model output is a real, detectable pattern; it's the edge cases (heavily edited text, non-native writing, short samples) where accuracy drops off.

What this means if you write with AI assistance

If you use AI tools as part of your writing process — for drafting, brainstorming, or getting past a blank page — the practical takeaway isn't about tricking a detector. It's that genuinely edited writing, with real variation in rhythm and specific, concrete detail, naturally moves away from the statistical patterns detectors are built to catch. That's a side effect of writing that's actually good, not a workaround.

Our Text Humanizer is built around that same principle: it varies sentence rhythm and removes the hedging, templated phrasing that makes text statistically uniform in the first place — the same qualities that make writing read naturally to a person also happen to move it away from a detector's core signals. For the specific patterns worth fixing, see why AI writing sounds robotic and AI vs. human writing: the real differences.

The detector landscape, and why they disagree

There is no single detection method. What gets marketed as "AI detection" covers several different approaches with different failure modes, which is a large part of why running the same passage through three tools returns three answers.

Perplexity-based classifiers are the approach described above. They run your text through a language model, measure how surprised it is, and threshold the result. Most free detectors work this way. They are cheap to build and inherit every weakness of the proxy.

Fine-tuned classifiers are trained directly on labelled datasets of human and machine text, learning whatever features separate the two in that dataset. These can outperform perplexity thresholds on text resembling their training data, and degrade sharply outside it. A classifier trained largely on student essays and GPT-3.5 output will behave unpredictably on technical documentation or a newer model's output.

Watermarking is the only approach with a real theoretical basis. The generating model biases its word choices according to a secret key, embedding a statistical signature that a party holding the key can verify. It works — but it requires cooperation from the model provider, only detects that specific provider's output, and is defeated by paraphrasing. It cannot help with the general question of whether arbitrary text was machine-written.

Metadata and provenance approaches sidestep detection entirely by attaching signed information about how content was created. C2PA and similar standards do this for images. This is the most promising direction for the long run, and it says nothing about text that carries no provenance data — which is nearly all text.

Most commercial detectors combine the first two. Their disagreements are not a bug being ironed out; they reflect genuinely different models measuring genuinely different things.

Why accuracy claims are hard to believe

Vendor accuracy figures — 99%, 98%, "less than 1% false positives" — should be read with three questions in mind.

Measured on what? Benchmark datasets typically pair clean human text with unedited model output. Real submissions are messier: partially edited, second-language, heavily proofread, or written by someone imitating a formal register. Accuracy on a clean benchmark tells you little about accuracy on a real cohort.

At what threshold? Any classifier trades false positives against false negatives. Quote a number without the operating threshold and you can make either look excellent. A detector tuned to catch nearly all machine text will flag far more innocent writing than the headline suggests.

Against which generation? Detectors are benchmarked against the models available when the test ran. Each new generation produces text closer to human statistical distributions, so measured accuracy decays over time even with no change to the detector.

Then there is the base rate. A detector with a genuine 1% false-positive rate, applied to 10,000 submissions of which 500 are actually machine-written, produces roughly 95 false accusations. The tool is working as specified. Ninety-five people are still wrongly accused.

The theoretical problem

Beyond engineering, there is a reason to doubt this ever gets solved.

Detection depends on machine text being statistically distinguishable from human text. But models are trained to imitate human text, and each generation does it better — the distributions converge by design. Work in this area has argued that as the two become sufficiently similar, no detector can do much better than chance without an unacceptable false-positive rate. Improving the model degrades the detector.

Paraphrasing compounds it. Running machine output through a rewriting pass disrupts the statistical signature detectors rely on, which is why they can be defeated by simple means. That cuts both ways, and it is why we are explicit that our own Text Humanizer is not a detection-evasion product: the same operation that makes a draft read better also erodes the signal, and treating that as a feature invites exactly the use we do not want.

What to do with a score

For anyone acting on a detector result:

Treat it as a prompt, not a finding. Turnitin's own guidance says its indicator should not be the sole basis for an academic integrity decision. A high score is a reason to ask a question, not to reach a conclusion.

Look at process evidence instead. Version history, drafts, notes, and a conversation about the work are all more informative than a percentage. Someone who wrote a piece can say why the third section exists; someone who generated it cannot.

Weight the population. If the writer is a non-native English speaker, the score means substantially less — the disparity is well documented, and we cover the mechanism in why AI detectors fail non-native English speakers.

Do not act on short samples. Below a few hundred words, these tools are close to guessing.

If you are on the receiving end of an accusation, our guide to defending your work covers what to preserve and what to say.

Common questions

What is perplexity in AI detection?
Perplexity measures how surprised a language model is by each word in a text. A model reads your sentence one word at a time and, at each position, checks how probable the actual word was. Low average surprise means low perplexity, which detectors read as machine-written, because models generate by repeatedly picking high-probability words. Human writing tends to score higher because people make less predictable choices.
What is burstiness in writing?
Burstiness describes how much sentence length and complexity vary across a document. Human writing lurches: a long winding sentence, then a short one, then a fragment. Generated text is more even, because a model optimises sentence by sentence with no sense of the shape of a page. Low burstiness is the second main signal detectors combine with perplexity.
Can AI detectors be wrong?
Frequently, in both directions. Even a genuine 1% false positive rate applied across 10,000 submissions wrongly flags around 95 people. Detectors also disagree with each other on the same passage, which shows they are estimating rather than measuring, and they are far less reliable on short samples, formal registers, and second-language writing.
Which AI detector is the most accurate?
No detector is reliably accurate enough to act on alone, and rankings between them shift with every model release. More usefully: they all rest on the same predictability proxy, so they share the same failure modes. If two detectors disagree sharply on the same passage, that is evidence about the tools rather than about the text.
Will AI detection get more accurate over time?
There are structural reasons to doubt it. Models are trained to imitate human writing, so the two statistical distributions converge as models improve, meaning better models make detection harder rather than easier. Work in this area has argued that beyond a point no detector can do much better than chance without an unacceptable false positive rate.
Do AI detectors work on paraphrased or edited text?
Much less well. Paraphrasing disrupts the statistical signature detectors rely on, which is why they can be defeated by simple means. This cuts both ways: it also means genuinely edited human writing can score very differently from the same author's first draft, which is part of why the tools are unreliable as evidence.
What is watermarking, and is it better than detection?
Watermarking biases a model's word choices according to a secret key, embedding a signature that whoever holds the key can verify. It is the only approach with a real theoretical basis, but it requires cooperation from the model provider, only detects that provider's output, and is defeated by paraphrasing. It cannot answer the general question of whether arbitrary text was machine-written.

Keep reading