How Do AI Detectors Work? A Complete Guide

Oct 02, 2026•20 min read
•Krish•AI
How Do AI Detectors Work? A Complete Guide

Quick Answer: How Do AI Detectors Work?

How do AI detectors work? Most of them look for statistical patterns that separate machine-written text from human-written text, then turn what they find into a probability. There are four main approaches:

  1. Trained classifiers. A model is trained on large collections of human-written and AI-written text until it can guess which is which. Most current commercial detectors work this way.
  2. Statistical and "zero-shot" tests. These measure how predictable a passage is to a language model, using ideas such as perplexity, probability curvature or the disagreement between two models. They need no labeled training set.
  3. Watermarks. The AI system itself nudges its word choices in a hidden, statistically detectable pattern, and a matching detector checks for it. This only works for text from systems that add the watermark.
  4. Retrieval and provenance. A provider keeps a record of what it generated and checks new text against it, or content carries signed metadata that says where it came from.

All of these produce evidence, not proof. They make mistakes in both directions, they get worse on short, edited or unusual text, and the research record shows they can be fooled by paraphrasing. The rest of this guide explains each method, what happens inside a typical detector, why errors occur, and how to read a score sensibly.


Introduction

When someone pastes an essay into a detector and sees "87% AI," it can feel like the software has found a fingerprint. It hasn't. What it has produced is a prediction from a statistical model, built on assumptions about how machine-written and human-written text differ.

Understanding those assumptions matters for three kinds of readers. Teachers and editors need to know how much weight a score deserves. Writers need to understand why their own work might be flagged. Developers and buyers need to evaluate detector claims, and the evidence shows that is worth doing carefully: in 2025 the US Federal Trade Commission proposed an order against a detection company over a claim of "98 percent" accuracy, saying independent testing found 53 percent on general-purpose content and that the model had only been trained to classify academic content (FTC).

This article explains how AI detection works in plain language, citing the research papers and vendor documentation behind each idea. It focuses on text detection, since that is what most people mean, and briefly covers images and other media. For a comparison of specific tools and their pricing, see our companion guide to the best AI detectors.


How AI Detectors Work: Language Models Leave Statistical Traces

To see why detection is possible at all, it helps to know how a language model writes. It generates text one small piece at a time, called a token. At each step it calculates a probability for every possible next token and then picks one, usually favoring likely choices, with some randomness. Human writers aren't doing that, and over many words the two processes can produce slightly different statistical profiles.

Detectors try to measure that difference. The signals include:

  • Predictability. Machine text tends to use words that a language model finds likely. Human text can include more surprising choices.
  • Variation. The predictability of human text may fluctuate more from sentence to sentence.
  • Vocabulary and structure patterns, such as favored phrases, transitions and sentence shapes.
  • Hidden signals, when a watermark has been deliberately added.

The catch is that these traces are weak and shrinking. Language models are trained to write like people, newer models are better at it, and the people who use them edit, mix and paraphrase the output. Researchers have argued that as models improve, the distributional gap between human and machine text narrows, which limits how reliable any detector can become (Sadasivan et al., arXiv).


Method 1: Perplexity, Burstiness and Token Statistics

Perplexity

Perplexity is a measure of how surprised a language model is by a piece of text. If the model finds each next word easy to predict, perplexity is low. If the words are unexpected, perplexity is high. Because AI systems tend to produce words that language models consider likely, low perplexity became an early signal of machine writing.

Burstiness

Burstiness refers to variation, for example in sentence length or in how predictable different sentences are. The idea is that people alternate between plain and unusual phrasing and between short and long sentences, while machine output is more even. Early detectors such as GPTZero were widely described as combining perplexity and burstiness. Reports say GPTZero later moved away from relying on those two measures to an end-to-end trained classifier, and its current documentation describes a trained classifier with sentence-level and document-level results (GPTZero FAQ).

GLTR: making the statistics visible

An influential early tool, GLTR, applied statistical methods to the artifacts that text generation models leave behind and showed them to readers as visual annotations. According to its authors, the annotation scheme improved the human detection rate of fake text from 54% to 72% without any prior training (Gehrmann, Strobelt and Rush, arXiv). It illustrates the underlying idea: how highly ranked each word is among the model's predictions carries information about whether a model or a person chose it.

Limits of simple statistics

Raw perplexity is a blunt tool. It depends on which language model is doing the measuring, and it can be affected by topic, genre and writing level. Formulaic human writing, such as templates, boilerplate or very plain prose, can have low perplexity. Creative AI output with high randomness can look "human." This is one reason most commercial tools moved on to trained classifiers.


Method 2: Trained Classifiers

How they are built

A trained classifier is a machine-learning model, often a neural network built on the same transformer architecture as language models. Developers build it in roughly these steps:

  1. Collect human-written text from many sources and genres.
  2. Generate AI text using many different language models, prompts and settings.
  3. Label each sample as human or AI.
  4. Train the model to predict the label from the text.
  5. Test it on held-out samples and tune a threshold for what counts as a flag.
  6. Update it as new language models appear.

GPTZero's documentation describes its detector as a classifier trained on a large, diverse corpus of human-written and AI-generated text, offering sentence-level and document-level analysis with confidence scores and highlighting (GPTZero FAQ). Pangram's technical report describes a transformer-based neural network trained to distinguish text written by large language models from text written by humans, using a technique it calls hard negative mining with synthetic mirrors (Emi and Spero, arXiv). That report is from the vendor, not an independent evaluation, so treat its performance claims as claims.

What the output means

A classifier outputs a probability, and the vendor chooses how to present it. GPTZero's FAQ explains that its percentages describe the detector's reliability on similar documents: "90% means that 90% of the time on similar documents our detector is correct in the prediction it makes." In other words, a score is not the share of the text that was written by AI. Always read the vendor's own explanation, because different tools use the same-looking number to mean different things.

Sentence-level and document-level scores

Many detectors score individual sentences and then combine them into a document-level result. GPTZero notes that the accuracy of its model increases as more text is submitted, so document-level results are more reliable than paragraph- or sentence-level ones. That is an important point for anyone tempted to act on a single highlighted sentence.

The weak point: the training data

A classifier is only as good as the examples it learned from. It can perform well on text that resembles its training data and poorly on text that doesn't. The Workado case is a stark example: according to the FTC, the model had only been trained or fine-tuned to classify academic content, and independent testing found it was accurate 53 percent of the time on general-purpose content. GPTZero's own FAQ acknowledges that the majority of its dataset is English prose written by adults, and that it can sometimes flag other machine-generated or highly procedural text as AI-generated. Detectors can therefore be unreliable on genres, languages and writers that are underrepresented in training.


Method 3: Zero-Shot and Model-Based Tests

Some detection methods don't train a dedicated classifier. Instead they use language models directly and look for a statistical signature.

DetectGPT and probability curvature

DetectGPT is based on the observation that text generated by a language model tends to sit in regions where small changes to the wording make the model's own probability go down, which the authors describe as negative curvature of the log probability function. The method perturbs a passage using another language model and compares probabilities, and it needs only the log probabilities of the model of interest and random perturbations of the passage. The authors reported improving detection of fake news articles generated by a 20-billion-parameter model from 0.81 AUROC for the strongest zero-shot baseline to 0.95 AUROC (Mitchell et al., arXiv). AUROC is a standard measure of how well a method separates two classes, where 0.5 is chance and 1.0 is perfect.

Binoculars: contrasting two models

The Binoculars method scores text by contrasting two closely related language models. Its authors report that this requires no training data and detects over 90% of generated samples from ChatGPT and other models at a false positive rate of 0.01% (Hans et al., arXiv). These are results from the authors' own evaluations, and real-world performance depends on the text and on how it has been edited.

Why research results and products differ

Zero-shot methods are attractive because they don't need labeled examples, but they depend on having access to suitable language models and can be sensitive to how the text was generated. Research papers report results under specific conditions. The RAID benchmark, a large shared test with over 6 million generations spanning 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies, evaluated 8 open-source and 4 closed-source detectors and concluded that current detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties and unseen generative models (Dugan et al., arXiv). A method that looks excellent on a clean test set can look much weaker on varied, realistic text.


Method 4: Watermarking

The basic idea

Instead of guessing, the AI developer can build a signal into the text as it is generated. A well-known scheme proposes selecting a randomized set of "green" tokens before each word is generated, then softly promoting green tokens during sampling. A detector that knows the scheme can then check whether a passage contains more green tokens than chance would predict, using a statistical test with interpretable p-values, without needing access to the model itself (Kirchenbauer et al., arXiv). The authors report that the watermark can be embedded with negligible impact on text quality.

SynthID-Text in practice

Google DeepMind published a production-oriented text watermarking scheme, SynthID-Text, in Nature in 2024. According to coverage of the work, it modifies only the sampling procedure and does not affect model training, detection is computationally efficient without using the underlying model, a live experiment assessed feedback from nearly 20 million Gemini responses, and Google deployed it in Gemini and later open-sourced it (Nature, summarized by TechXplore).

Limits of watermarks

  • They only work on text from systems that apply them. A watermark says nothing about text produced by a model that doesn't add one.
  • They can be weakened. Researchers have shown that attackers can infer hidden signatures without access to the detection method, and that paraphrasing can reduce detection (Sadasivan et al.).
  • They need short-text care. A statistical test needs enough words to be confident.
  • They raise standards questions. Detection requires coordination between the generator and the checker.

Most commercial "AI detectors" that you paste text into do not rely on watermarks, because most text from most tools carries none.


Method 5: Retrieval and Provenance

Retrieval-based detection

One proposed defense against paraphrasing is for a provider to keep a database of sequences it has generated and check new text against it. Researchers who built a strong paraphrasing model reported that their retrieval-based defense detected 80% to 97% of paraphrased generations while keeping a 1% false positive rate on human writing (Krishna et al., arXiv). The approach requires the provider to store its outputs, which has privacy and scale implications, and it only covers text from that provider.

Content provenance

For images, audio and video, provenance standards offer an alternative to guessing. The C2PA's Content Credentials is described as an open technical standard for publishers, creators and consumers to establish the origin and edits of digital content (C2PA). Provenance records what a trusted tool says happened to a file, which is different from inferring it from the content. It helps only where the signal is present and preserved.


How Does AI Detection Work Step by Step? Inside a Typical Detector

Putting the pieces together, a detector that you paste text into usually does something like this.

  1. Receives the text and cleans it, for example removing formatting.
  2. Splits it into tokens and often into sentences or segments.
  3. Checks the length. Very short text may be rejected or flagged as unreliable.
  4. Computes features or runs the model. This might be a trained neural classifier, a perplexity-style statistic, or both.
  5. Scores each segment to estimate how likely it is to be machine-generated.
  6. Aggregates the segment scores into a document-level result.
  7. Applies a threshold that decides when to call something AI-generated, and may add a band such as "uncertain."
  8. Displays the result, often with highlighting.

Every step involves design choices, such as how long a segment must be, where the threshold sits and how the number is labeled. Two detectors can therefore disagree about the same text without either being "broken," and this is why a single score should not be the end of an inquiry.


Why Detectors Get It Wrong

Understanding the mechanics explains most of the failure modes.

Short text. There is little evidence in a few sentences. A 2025 study found that commercial detectors generally lost some accuracy on passages under 50 words (Chicago Booth Review summary of Jabarian and Imas).

Paraphrasing and editing. If a person rewrites machine text, or software paraphrases it, the statistical signature changes. The researchers behind DIPPER, an 11-billion-parameter paraphrasing model, reported reducing DetectGPT's detection accuracy from 70.3% to 4.6% while keeping the meaning, and showed it also evaded watermarking, GPTZero and OpenAI's classifier (Krishna et al.). Those results are from 2023, and detectors have changed, but the principle holds.

Training-data mismatch. A detector trained on one kind of text can fail on another, as the Workado case shows.

Bias against some writers. Researchers found that GPT detectors consistently misclassified non-native English writing samples as AI-generated while accurately identifying native writing, and warned against using them in evaluative or educational settings (Liang et al., arXiv). A plausible reason, suggested by the findings, is that writing with a more constrained vocabulary and grammar can look statistically similar to machine output. Formulaic, highly procedural human text can be flagged for similar reasons, as GPTZero's FAQ acknowledges.

Unseen models and settings. The RAID benchmark found detectors can be fooled by variations in sampling strategies and by generative models they have not seen.

Mixed authorship. A person may draft, use an assistant to fix grammar, and rewrite. The text is neither purely human nor purely machine, and detectors are generally built for the extremes.

Theoretical limits. As language models become better at imitating people, the statistical gap shrinks, and some researchers argue this places inherent limits on reliable detection.


The Base Rate Problem: Why a Small Error Rate Still Hurts

Even an accurate-sounding detector can mislead when you use it on a population in which AI use is uncommon. This is a hypothetical illustration of the arithmetic, not data about a real tool.

Suppose a school checks 1,000 essays. Assume 5% were written mostly by AI, so 50 essays are AI and 950 are human. Assume the detector catches 95% of AI essays (a sensitivity of 95%) and wrongly flags 1% of human essays (a false positive rate of 1%).

  • It correctly flags about 47.5 AI essays (95% of 50).
  • It wrongly flags about 9.5 human essays (1% of 950).
  • Of the roughly 57 essays flagged, about 17% are actually human.

If AI use were much more common, the picture would change. With 20% AI essays (200 AI, 800 human), the detector would flag about 190 AI essays and 8 human ones, so only about 4% of flagged essays would be human. If the false positive rate were 0.1% instead of 1% in the first scenario, about one human essay would be flagged out of roughly 48 flags, or about 2%.

The lesson is that how many flagged people are innocent depends on three things: the detector's false positive rate, its sensitivity and how common AI use is in the group you're checking. This is the reasoning behind Vanderbilt's decision to disable Turnitin's AI detector in 2023: it calculated that a claimed 1% false positive rate could still have incorrectly flagged around 750 of 75,000 papers a year (Vanderbilt).


How Detectors Are Evaluated

If you read a detector's claims or a research paper, a few terms explain what the numbers mean.

TermMeaning
False positive rate (FPR)The share of human-written text wrongly flagged as AI
False negative rate (FNR)The share of AI-written text wrongly called human
Sensitivity (recall)The share of AI text correctly caught, equal to one minus the FNR
PrecisionOf the texts flagged as AI, the share that really are AI
AUROCA summary of how well a method separates two classes across all thresholds, where 0.5 is chance
ThresholdThe score above which a text is called AI, which trades false positives against false negatives

Which error matters more depends on the setting. In an academic integrity case, a false positive can harm an innocent student, so people often want a very low FPR. A 2025 NBER study introduced the idea of a "policy cap," a limit on tolerable false positives or negatives, and reported that one detector was the only tool to satisfy a strict cap (FPR at or below 0.005) without sacrificing accuracy (Jabarian and Imas). That study also found all three commercial tools it tested kept false positive rates below 1% in its dataset. Results like these depend on the dataset, so test a detector on text like yours.


How "Humanizers" Work, and Why They Matter

"Humanizer" tools rewrite AI text so that detectors score it as human. Conceptually, they work by changing the statistical features that detectors rely on, for example by paraphrasing, varying sentence structure and swapping words. The research on paraphrasing attacks shows why this can succeed against some detectors: if the signal is a pattern in word choice, rewriting the words changes the pattern.

This creates an arms race. Detector vendors retrain against known humanizers, and humanizer vendors adapt. A 2025 study reported that its best-performing detector remained robust against humanizer tools, but also cautioned that performance will likely vary as the field evolves (Chicago Booth Review). For anyone relying on detection, the practical point is that a "pass" does not prove a human wrote something, and a flag does not prove a machine did.


How to Read a Detector Result

Because of all of the above, a sensible reading of a result is cautious.

  1. Check what the number means. Read the vendor's explanation. A "percent AI" score may describe confidence, not the share of text.
  2. Check the length. Short passages produce weak evidence.
  3. Look at the pattern, not one sentence. Highlights on a few sentences are weak evidence on their own.
  4. Consider the writer and the genre. Non-native writers and formulaic genres are at higher risk of false flags.
  5. Compare more than one tool, and treat disagreement as information.
  6. Look for other evidence: drafts, version history, notes, earlier writing and the writer's ability to explain the work.
  7. Give the writer a chance to respond before acting.

GPTZero's FAQ itself says results should not be used to punish students and that detection should be used as part of a holistic assessment. For more on using detectors fairly and choosing a tool, see our guide to the best AI detectors, and for setting organizational rules around AI use, see what AI governance is.


How to Judge a Detector's Claims

When a vendor says its detector is accurate, ask:

  • Accurate on what? Which genres, lengths, languages and models were in the test?
  • What were the false positive and false negative rates, not just a single accuracy figure?
  • Was the evaluation independent, or run by the vendor?
  • How does it handle edited, paraphrased and short text?
  • How is the score defined, and what does the vendor advise you to do with it?
  • How often is it updated as new models appear?
  • What happens to submitted text, including storage and training use?

Our buyer's checklist for AI-powered products offers a wider framework for separating real capability from marketing.


Where AI Detection Is Heading

Several trends are visible, though none is settled.

  • Provenance over inference. Standards such as Content Credentials aim to record where media came from, instead of guessing afterward.
  • Watermarking at scale. Deployments such as SynthID-Text show it is technically feasible, but adoption across providers, and resistance to paraphrasing, remain open questions.
  • Better trained classifiers, especially for specific domains, with continuing competition between detectors and humanizers.
  • More emphasis on process, such as drafts, version history and oral follow-up, because they don't depend on a classifier being right.
  • Regulation and standards. Rules about disclosure, labeling and claims, as in the FTC action, are likely to shape how detectors are marketed.

The most robust practices combine technology with human judgment and clear policies.


Glossary

  • Token: a small unit of text, such as a word or part of a word, that a language model processes.
  • Perplexity: a measure of how surprised a language model is by a text.
  • Burstiness: variation in sentence length or predictability across a text.
  • Classifier: a model trained to assign labels, such as human or AI.
  • Zero-shot detection: detection that uses language models directly without training a dedicated classifier.
  • Watermark: a hidden statistical pattern added to generated text so that it can be identified later.
  • Provenance: a record of where content came from and how it was edited.
  • False positive: human text flagged as AI.
  • False negative: AI text called human.
  • Threshold: the cut-off score that determines a flag.
  • Humanizer: a tool that rewrites AI text to change how detectors score it.

Frequently Asked Questions

How do AI detectors work? Most analyze a text with a model trained on examples of human and AI writing, or with statistical tests based on how predictable the text is to a language model. They output a probability, often with sentence-level highlights. Some research tools use watermarks or retrieval instead. None of these methods produces proof.

Do AI detectors use perplexity and burstiness? Early detectors were widely described as using them. Perplexity measures how predictable a text is to a language model, and burstiness measures variation. Reports say GPTZero later moved to a trained classifier, and its current documentation describes a classifier with sentence and document-level results.

What does a "90% AI" score mean? It depends on the vendor. GPTZero's FAQ explains that a 90% figure means the detector is correct about 90% of the time on similar documents, not that 90% of the text is AI-written. Always read the tool's definition.

Can AI detectors tell which AI wrote a text? Generally no. Most detectors estimate whether a text is machine-generated, not which model produced it. Watermark detection is tied to the specific scheme that added the watermark.

Can AI detectors detect paraphrased text? Often poorly. Research showed a paraphrasing model could sharply reduce the accuracy of several detectors, and a retrieval-based defense could catch many paraphrased generations but requires a provider to store what it generated.

Is watermarking the same as AI detection? No. Watermarking adds a hidden signal when text is generated, so it only works for text from systems that add it. Typical detectors you paste text into infer from statistical patterns and don't rely on a watermark.

Why do AI detectors flag human writing? Human writing can resemble machine output statistically, especially if it is formulaic, simple or written by someone with a constrained vocabulary. Research found detectors misclassify non-native English writing at higher rates, and detectors can also be thrown off by genres they were not trained on.

How accurate are AI detectors? It varies by tool, text and date. A 2023 study found tools "neither accurate nor reliable," while a 2025 study found leading commercial tools performed much better on its dataset. See our guide to the best AI detectors for the evidence.

Do AI detectors store the text I submit? It depends on the vendor. Check each tool's privacy policy for storage, retention and whether text is used to train models, especially for confidential or student work.

Can a detector prove someone used AI? No. A detector's output is a probability based on patterns. It should be one input alongside drafts, version history and a conversation with the writer.

Will AI detection get better or worse? Both directions are plausible. Detectors are improving, but language models are also getting better at imitating people, and tools for evading detection are improving. Some researchers argue there are theoretical limits to reliable detection as models advance.

Do detectors work on images and audio? Different methods apply. Some systems analyze content for artifacts, and provenance standards such as Content Credentials record the origin and edits of media. Reliability varies, and provenance only helps where the signal is present.


Conclusion

How do AI detectors work? They estimate, from statistical patterns, whether text looks more like a language model's output or a person's. Trained classifiers learn the difference from examples, perplexity-style and zero-shot tests measure how predictable the text is, watermarks embed a deliberate signal, and retrieval and provenance methods rely on records instead of inference.

Each approach has limits that follow directly from how it works. Short text is thin evidence, paraphrasing and mixed authorship blur the signal, training data shapes what a classifier can recognize, some writers are flagged more often than others, and even a small false positive rate produces many wrongly flagged people at scale. The sensible way to use a detector is as one input: understand what its score means, test it on text like yours, look for other evidence, and give the writer a fair chance to respond.


Sources

Tags

#AI Detectors#AI Detection#Machine Learning#Watermarking#Academic Integrity#Generative AI