
Quick Answer: What Is the Best AI Detector?
On the best independent evidence available, Pangram came out strongest in a 2025 academic evaluation of commercial detectors: it had the lowest false positive rate and was the only tool to meet a strict cap on false positives. GPTZero also did well, and in one summary it missed the least AI text, while Originality.ai missed noticeably more. But that is one study, on one set of texts, and the broader research record is cautious: earlier independent tests found that detection tools were neither accurate nor reliable, and detectors have been shown to misjudge writing by non-native English speakers.
So the best AI detector is the one you use as a screening signal, never as proof. The practical rules are:
- Pick a detector with independent evidence behind it, not only its own accuracy claims.
- Test it on your own kind of text before you rely on it.
- Never treat a score as the sole basis for an accusation, a grade or a rejection. Combine it with drafts, version history, a conversation with the writer and editorial judgment.
- Expect errors in both directions: human writing flagged as AI, and AI writing that slips through, especially if it has been edited or run through a "humanizer."
This guide compares eight of the best AI detectors available, summarizes what independent researchers have found, shows current pricing where I could read it, and explains how to use a detector responsibly.
Introduction
AI writing tools have created a new demand: telling human text from machine text at scale. Teachers want to know whether an essay was written by a student, editors want to know whether a submission was generated, and content teams want to check freelance work. A market of detectors, often sold as an AI content detector, has grown up to meet that demand, and most promise high accuracy.
The difficulty is that accuracy claims are easy to make and hard to verify. In April 2025 the US Federal Trade Commission announced a proposed order against a detection company, Workado, over its claim that its AI Content Detector was "98 percent" accurate. The FTC said independent testing showed accuracy on general-purpose content of just 53 percent, and that the model had only been trained or fine-tuned to classify academic content (FTC press release). Press reports say the order was later finalized. Fifty-three percent is little better than a coin toss. That case does not describe every detector, but it is a good reason to ask any vendor what its numbers were measured on.
This article is not based on tests we ran ourselves. It brings together independent research, vendors' published pricing and features, and clear guidance on how to evaluate a detector yourself. Where a figure comes from a vendor or from a news summary instead of the primary source, the article says so, and every outside figure is linked.
How AI Detectors Work
Most detectors are classifiers. They are trained on large collections of human-written and machine-written text and learn statistical differences between them, then output a probability or a score for a new piece of text. Some also highlight individual sentences. Vendors rarely publish full details of their training data or methods, which is one of the criticisms that universities have raised.
Several things follow from that design:
- A score is a probability, not a fact. "92% AI" means the model is confident under its own assumptions, not that 92% of the text was machine-written.
- Performance depends on the training data. A detector trained mostly on essays may behave differently on emails, résumés, technical writing or poetry.
- Text length matters. In one 2025 study, the commercial tools generally lost some accuracy on passages under 50 words.
- Editing changes the picture. Paraphrasing, mixing human and AI writing, and "humanizer" tools can all move a score.
- Models keep changing. A detector built against one generation of language models can lose accuracy when new models appear.
AI Detector Accuracy: What Independent Research Shows
The most useful evidence comes from researchers with no commercial stake. Here is what the main studies say.
2023: "Neither accurate nor reliable"
A research group led by Debora Weber-Wulff tested 14 detection systems, including Turnitin and PlagiarismCheck alongside 12 publicly available tools. They concluded that the available tools were "neither accurate nor reliable," found a main bias toward classifying output as human-written instead of detecting AI-generated text, and reported that content obfuscation techniques significantly worsened performance (Weber-Wulff et al., arXiv). Those tools and models have changed since, but the study remains a useful baseline.
2023: Bias against non-native English writers
Weixin Liang, James Zou and colleagues found that GPT detectors consistently misclassified non-native English writing samples as AI-generated, while native writing samples were accurately identified. They warned against using detectors in evaluative or educational settings, particularly where they could penalize non-native English speakers (Liang et al., arXiv). They also found that simple prompting strategies could mitigate the bias and could also bypass the detectors, which shows how brittle the signals are.
2023: OpenAI withdrew its own detector
OpenAI released an AI text classifier in January 2023 and withdrew it in July 2023 because of its "low rate of accuracy" (Search Engine Land). OpenAI's own announcement reported that on its test set the classifier correctly identified about 26% of AI-written text and incorrectly labeled about 9% of human-written text as AI-written (OpenAI). I could not open that page directly, so these two figures are as widely reported and should be checked at the source. The episode shows that even a leading AI developer withdrew its detector for accuracy reasons.
2023: Vanderbilt disables Turnitin's AI detector
On August 16, 2023, Vanderbilt University announced it was disabling Turnitin's AI detection feature. Its reasons included that a claimed 1% false positive rate could still have incorrectly flagged around 750 of 75,000 papers submitted in a year, that Turnitin gave no detailed information on how it decides, and that detectors have been found more likely to label non-native English writing as AI-written. Vanderbilt concluded it did not believe AI detection software was an effective tool that should be used (Vanderbilt). The arithmetic is worth remembering: a small false positive rate becomes a large number of wrongly flagged people at scale.
2025: A more favorable result for the best commercial tools
Brian Jabarian and Alex Imas evaluated three commercial detectors, Pangram, OriginalityAI and GPTZero, plus an open-source RoBERTa classifier, on a large corpus. Their NBER working paper reports that commercial detectors outperformed the open-source one, and that Pangram achieved near-zero false negative and false positive rates that remained robust across models, threshold rules, very short passages and "humanizer" tools (NBER working paper, September 2025). A Chicago Booth summary of the work adds detail: the dataset had about 2,000 human-written passages across six types of writing, with AI versions generated by four language models, and all three commercial tools kept false positive rates below 1%. False negative rates were roughly 0% to 2% for GPTZero, 2% to 4% for Pangram and 10% to 40% for Originality.ai, and all the commercial tools lost some accuracy on passages under 50 words, although the paper reports that Pangram's rates stayed near zero even on such short "stubs" (Chicago Booth Review).
The authors also caution that detection is a fast-moving field in which performance will likely change as detectors and humanizers compete.
What to take from the research
- Results vary a great deal by tool and by date. Early tests were poor, and a later study found much better performance from the leading commercial tools.
- Even good tools make mistakes, and the cost of a false positive falls on a real person.
- Bias risks are documented, especially for non-native English writers.
- Vendor accuracy claims should be backed by tests on text like yours.
- Short text and edited text are the hard cases.
Best AI Detectors: Comparison Table
"Vendor" means I read the price on the vendor's own page in October 2026. Where a page didn't show prices, the table says to check the vendor. The evidence column records whether I found independent evidence on the tool, not a ranking.
| Tool | Best known for | Free option | Paid pricing (USD) | Independent evidence reviewed |
|---|---|---|---|---|
| Pangram | AI detection with API and plagiarism checks | Yes, 2,000 words a day | Individual $20, Professional $65 a month | 2025 NBER study: near-zero error rates |
| GPTZero | Education-focused detection | Check vendor | Check vendor | 2025 NBER study: low false positives |
| Originality.ai | Detection plus plagiarism and site scans for publishers | Limited, 3 scans a day | Pro $12.95 a month (annual) | 2025 NBER study: higher false negatives |
| Winston AI | Detection with plagiarism and image checks | 14-day trial | Essential $18, Advanced $29, Elite $49 a month | None reviewed |
| Copyleaks | Detection with plagiarism, institutional focus | Check vendor | Check vendor | None reviewed |
| Sapling | Detector inside a writing assistant, with an API | Yes, 2,000 characters per check | Pro $25 a month | None reviewed |
| Turnitin | Institutional detection inside education tools | No, sold to institutions | Institutional | Tested in a 2023 study; disabled by some universities |
| Open-source classifiers | Free research-style detectors | Yes | Free | 2025 study found one performed close to random guessing |
The Tools in More Detail
1. Pangram
Pangram offers an AI detector with a free tier, a browser extension, Google Docs integration and an API. According to Pangram's pricing page, the free tier scans up to 2,000 words a day, the Individual plan is $20 a month with up to 300,000 words a month and plagiarism detection, and the Professional plan is $65 a month with five times the scanning allowance and $200 of monthly API credits. In the Jabarian and Imas study it was the only tool that met a strict false positive cap without losing accuracy. Remember that this is one study, and results on your own text may differ.
Best for: teams that want a strong independent track record and API access. Watch out for: performance on very short text, and the fact that no tool is perfect.
2. GPTZero
GPTZero is widely used in education and was one of three commercial tools in the 2025 study, where its false positive rate was below 1% and its false negative rate was low. Its pricing page did not show prices when I read it, so check GPTZero's site for current plans and word limits.
Best for: educators and writers who want a tool with independent evidence behind it. Watch out for: how its results are presented in an academic integrity process, since a score is not proof.
3. Originality.ai
Originality.ai combines AI detection with plagiarism checking, readability and grammar tools, fact checking and full-site scans, and is aimed at publishers and content teams. According to its pricing page, the free Basic plan gives 60 credits a day (up to three scans, AI detection only, 2,000 words per scan), Pro is $12.95 a month billed annually ($14.95 monthly) with 2,000 monthly credits, and Enterprise is $136.58 a month annually with API access and a longer scan history. One credit equals 100 words. In the 2025 study its false negative rate was higher than the other commercial tools, between 10% and 40%, meaning more AI text went undetected.
Best for: publishers who also want plagiarism and site-scanning tools in one place. Watch out for: missed AI text in the independent test, and credit limits on long documents.
4. Winston AI
Winston AI includes AI text detection, image and deepfake detection, a plagiarism checker and OCR. According to its pricing page, there is a free 14-day trial with 2,000 credits, then Essential at $18 a month ($9 on annual billing), Advanced at $29 ($14.50 annually) and Elite at $49 ($24.50 annually), with credit allowances that grow by tier. I did not find an independent accuracy evaluation of it, so test it on your own samples.
Best for: users who want text and image checks in one tool. Watch out for: limited independent evidence.
5. Copyleaks
Copyleaks is known for plagiarism detection with an AI detector aimed at education and enterprise. Its pricing page blocked my fetch, so I can't give verified prices. Check Copyleaks for current plans. I did not find an independent evaluation to cite.
Best for: organizations that already use it for plagiarism checks. Watch out for: verifying accuracy on your kind of content.
6. Sapling
Sapling's AI detector sits alongside its writing assistant and API. According to its pricing page, the free tier checks up to 2,000 characters at a time, Pro is $25 a month with 100,000 characters per check, and the API is usage-based with a $5 minimum. The free character limit makes it unsuitable for checking whole essays without paying.
Best for: quick checks of short passages and developers who want an API. Watch out for: short-text accuracy limits noted in the research, and no independent evaluation found.
7. Turnitin AI Detector
Turnitin's AI writing detection is sold to institutions inside its integrity tools, so individuals generally can't buy it directly. Turnitin has stated that its document-level false positive rate is below 1% for documents with more than 20% AI writing and that its sentence-level false positive rate is about 4%, and it advises educators to use the score to start a conversation with students, not to make a determination of misconduct (Turnitin). Turnitin's page was not readable to me, so these figures follow its published statements as reported in search results and should be checked at the source. Some universities, including Vanderbilt, have disabled the feature. Turnitin was included in the 2023 Weber-Wulff study.
Best for: institutions that already use Turnitin and have a clear policy for how scores are used. Watch out for: treating a score as evidence, especially for non-native English writers.
8. Open-source classifiers
Free, open-source classifiers such as RoBERTa-based models are available for researchers and developers. In the 2025 study, the open-source model's accuracy was close to that of random guessing, so they are poor choices for decisions about real people.
Best for: research and experimentation. Watch out for: using them for any decision that affects someone.
How to Choose a Detector
Match the tool to the job, and test it before you rely on it.
- Check the evidence. Is there independent testing, on text like yours, with false positive and false negative rates reported?
- Check the scope. Does it handle your languages, text lengths and genres?
- Check how results are shown. A single percentage invites over-interpretation. Sentence-level highlights and clear explanations are more useful.
- Check the data policy. Where is submitted text stored, and is it used for training? This matters for student work, confidential drafts and client material.
- Check the pricing model. Word or credit limits can make long documents expensive.
- Check the integration. Browser extensions, Google Docs add-ons and APIs fit different workflows.
- Check the vendor's claims. Ask what the accuracy number was measured on, and whether an independent party reproduced it.
For a broader approach to evaluating AI products and the claims made for them, see our buyer's checklist for AI-powered products.
How to Test a Detector Yourself
If you will use a detector regularly, a short test of your own is worth the time. It takes a few hours.
- Collect known human text. Gather 30 to 50 samples of writing you are certain was written by people, in the genres you care about. Include writing by non-native English speakers if they are among your users.
- Create known AI text. Generate matching samples with the AI tools your community is likely to use.
- Add the hard cases. Include AI text that a person has edited, human text that has been polished with grammar tools, and short passages.
- Run each detector on the same set, using the same settings.
- Count the errors. The false positive rate is the share of human text flagged as AI. The false negative rate is the share of AI text called human. Record both.
- Look at the pattern. Do errors cluster in certain genres, lengths or writers?
- Decide your tolerance. For decisions that affect people, the false positive rate matters most, and even a rate of 1% means many wrongly flagged people at scale, as the Vanderbilt calculation shows.
- Repeat periodically, because tools and models change.
This kind of test also shows you what a score really means in your setting.
Using AI Detectors Responsibly
If you use detection in a school, newsroom or company, these practices reduce harm.
- Use scores as a prompt for a conversation, not a verdict.
- Look for corroborating evidence: version history, drafts, notes, earlier writing by the same person and the writer's ability to explain the work.
- Give the writer a chance to respond before any action, and have an appeals process.
- Be careful with non-native English writers, given the documented bias.
- Don't rely on a single tool. Different detectors can disagree, and disagreement is itself information.
- Set a clear policy on what AI use is allowed, so that "AI-assisted" and "AI-generated" aren't confused. Grammar correction and brainstorming are different from submitting machine-written work.
- Protect privacy. Understand where the text goes when you paste it into a tool.
- Document decisions and the evidence behind them.
For guidance on setting organizational rules around AI use, see our guide to what AI governance is.
For Writers: If You Are Flagged Unfairly
AI detector false positives happen, and if a detector flags your own writing, these steps can help.
- Stay calm and ask what the score is based on. A score is not proof.
- Show your process: drafts, version history in your word processor, notes, outlines and sources.
- Offer to explain your work in person or in writing.
- Point to the limits of the tool, using the independent research above, including the documented bias against non-native English writers.
- Keep records going forward. Save drafts and use a tool that keeps version history.
- Ask for the policy. Find out how the institution or publisher treats detector results.
What About SEO and Content Marketing?
Content teams often ask whether detectors matter for search. Google has said it distinguishes between using AI as a tool and using automation primarily to manipulate rankings, and that it aims to reward helpful, reliable, people-first content (Google Search Central). In other words, the issue for search is quality, accuracy and usefulness, not whether a detector flags the text. A detector score doesn't tell you whether content is accurate or useful, and a human-written page can be thin or wrong. If you manage content, a better use of time is editorial review, fact-checking and adding original experience and sources. Our guide to generative engine optimization covers how content is found and cited in AI-driven search.
Alternatives to Detection
Because detection is imperfect, many educators and editors rely more on process.
- Assess the process, not only the product: drafts, outlines, annotated bibliographies, reflections.
- Use in-class or oral components where writers explain their reasoning.
- Ask for version history from the writing tool.
- Design assignments that depend on specific context, such as local data, class discussions or personal analysis.
- State clearly what AI use is permitted and how to disclose it.
- Build editorial standards for freelancers and contributors, including disclosure requirements.
These approaches are slower, but they don't depend on a classifier being right.
Common Mistakes
- Treating a score as proof.
- Trusting vendor accuracy claims without checking what they were tested on.
- Testing on only easy examples. Include short, edited and non-native English text.
- Ignoring false positives. They land on real people.
- Using one tool. Compare several, and don't treat a single flag as conclusive.
- Skipping the policy. Without clear rules, detectors create disputes.
- Pasting confidential text into a tool without checking its data policy.
- Assuming results stay the same. Tools and models change quickly.
Frequently Asked Questions
What is the best AI detector? There is no single best for every use. In a 2025 independent study, Pangram had the lowest false positive rate among three commercial tools and was the only one to meet a strict cap on false positives, GPTZero was also strong, and Originality.ai missed more AI text. Earlier research found detectors unreliable, so treat any score as a signal, test the tool on your own text, and never rely on it alone.
Are AI detectors accurate? It depends on the tool, the text and the date. A 2023 study found tools "neither accurate nor reliable," while a 2025 study found the leading commercial tools performed much better, with false positive rates below 1% in its dataset. Short text and edited text remain hard.
Can AI detectors be wrong? Yes, in both directions. They can flag human writing as AI and miss AI writing. Research has shown they can be biased against non-native English writers.
Is there a free AI detector? Several have free tiers with limits, including Pangram (2,000 words a day), Originality.ai (a limited daily allowance), Sapling (2,000 characters per check) and a 14-day trial for Winston AI. Check each vendor's current limits and data policy.
Does Turnitin's AI detector work? Turnitin states its document-level false positive rate is below 1% for documents with more than 20% AI writing and recommends using the score to start a conversation. Some universities have disabled the feature, citing concerns about false positives and transparency.
Can AI detectors detect ChatGPT or other specific models? Detectors look for statistical patterns, not a specific model's fingerprint. Their accuracy varies by model, and new models can reduce a detector's accuracy until it is updated.
Do AI humanizers beat detectors? Some do, and the research shows that editing and obfuscation can reduce detection performance. One 2025 study reported that its best-performing detector remained robust against humanizer tools, but results vary and the field changes quickly.
Are AI detectors biased? Research has found that detectors misclassify non-native English writing as AI-generated at higher rates than native writing, which is a reason for caution in schools and hiring.
Should teachers use AI detectors? Researchers cited above warn against using detectors in evaluative settings, and Turnitin itself advises using its score to start a conversation, not as a determination of misconduct. If used, a detector should be one input among several, with a fair process for students to respond.
Does Google penalize AI content? Google has said it focuses on whether content is helpful, reliable and people-first, and that automation is not prohibited in itself, while using it mainly to manipulate rankings is treated as spam. Read its current guidance, since policies change.
Is it legal to use an AI detector? Using one is generally legal, but how you use the result may be regulated, for example in education, employment and privacy law. Consider data protection and fairness rules, and take advice where decisions affect people.
How can I check whether a detector's accuracy claim is real? Ask what data it was tested on, whether the test was independent, and what the false positive and false negative rates were. The FTC's proposed order against Workado over a "98 percent" claim shows why this matters.
Conclusion
The best AI detector depends on what you need and how you will use it. On the independent evidence, a few commercial tools, led on false positives in one 2025 study by Pangram with GPTZero also performing well, can be accurate on the kinds of text that study used, while earlier research and real-world experience show serious limits: false positives, bias against non-native English writers, weak performance on short or edited text, and accuracy claims that can be misleading.
Use detectors as a screening signal. Test them on your own samples, understand their error rates, protect people from wrongful accusations with a fair process, and rely on drafts, version history and conversation for anything that matters. The goal is not to catch every machine-written sentence. It is to make good decisions about real people and real content.
Sources
- Weber-Wulff et al., Testing of Detection Tools for AI-Generated Text, arXiv (2023)
- Liang, Yuksekgonul, Mao, Wu and Zou, GPT detectors are biased against non-native English writers, arXiv (2023)
- Jabarian and Imas, Artificial Writing and Automated Detection, NBER Working Paper (September 2025)
- Chicago Booth Review, Do AI Detectors Work Well Enough to Trust?
- Vanderbilt University, Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector (August 2023)
- Federal Trade Commission, FTC Order Requires Workado to Back Up Artificial Intelligence Detection Claims (April 2025, proposed order)
- OpenAI, New AI classifier for indicating AI-written text (January 2023; page not readable by me)
- Search Engine Land, OpenAI's AI Text Classifier no longer available due to low rate of accuracy (2023)
- Turnitin, Understanding false positives within our AI writing detection capabilities
- Google Search Central, Google Search and AI-generated content
- Vendor pricing pages: Pangram, Originality.ai, Winston AI, Sapling; GPTZero and Copyleaks pricing was not readable and should be checked on their sites.
