Can AI Detectors Be Wrong?
Yes, and the how is also interesting. Detectors break in patterned and somewhat predictable ways which makes some writers’ lives a lot harder than others. Although headline accuracy is important, it is far less important than this.
Can AI Detectors Be Wrong: The Basics
Errors occur in both directions. A false positive means the system incorrectly flags writing as self-authored. A false negative means the system incorrectly flags writing as automatically generated. Vendors provide both rates, and the rates usually look good, but a rate measured over a balanced test set behaves very differently when dealing with a constant stream of real submissions where almost everything is human. When the honest population tends to be large, even a small false-positive rate will result in a constant stream of false positives.
There is evidence that this is more than just a hypothetical problem. In 2023, OpenAI removed their AI Text Classifier because they said it was not accurate enough. There have also been recent reports of detectors flagging long standing public domain documents as machine written. This sounds preposterous, but remember that this is a predictably mechanistic tool. Huge swaths of the Internet have been trained on texts that have been quoted and reproduced many times, and predictability is what this tool does best.
Published metrics on accuracy hide a tradeoff. They evaluate performance on the vendor’s internal assessment data. This will likely mean prompt output against clean human writing in the same genre. Your text is likely to look dissimilar. A translated case study, a heavily edited trade article, a specification sheet, or a draft that a person wrote and then tightened by an editor are all examples that fall beyond the distribution measured.
- That an ML algorithm wrote the text
- That the writer relied on any software
- That the text is erroneous, unoriginal, or plagiarized
- That a second detector would agree
- That the text would be evaluated the same way in the future
Why Can AI Detectors Be Wrong Matters
The two types of errors are priced very differently. A false negative means a client publishes a subpar output, which they notice at some point, and proceeds to commission a new one. A false positive risk includes a failed module, a disciplinary hearing, a terminated contract, or a freelancer being dropped silently and without reason from a roster. Same metrics, different impacts, and the impact befalls the one with the least amount of power in the transaction.
The bias is not equal. An analysis published in 2023 by Stanford researchers in Patterns performed several popular detection systems on some samples of TOEFL essays composed by non-native English speakers and on some samples of essays generated by students of US schools. More than half of the essays composed by non-native speakers were labeled as essays generated by AI, whereas the essays composed by native speakers were accurately classified almost every time. This behavior is explained by perplexity. A restricted lexicon and more predictable pattern in sentence generation lead to essays of low-surprise, the kind that is penalized by these generation detectors.
Writing that is bound by constraints other than the imagination of the writer will have similar effects. Technical writing, summaries of medical cases, drafting legal documents, regulated financial writing, and anything else that’s done with a template will have the same constraints. As will translation. So will good editing, because its primary goal is to remove all ambiguity and surprise. The better a copy editor does their job, the more machine-like the writing will score.
Do not make a decision on a score alone
When a result jeopardizes a person’s academic or professional standing, the grade, or the contract, the score represents the start of an investigation. The very first steps of the investigation consist of requesting the passages that were flagged, the tool and threshold applied, and the complete collection of the author’s drafts.
How Articled Approaches Can AI Detectors Be Wrong
Detector reports record the readings of twelve tools named by us, on the date of our report. They, in no way, certify anything. We do not guarantee that your detector will agree with ours, since we cannot dictate which tool you will use, which version you will use it on, and where its threshold will be this quarter.
If a draft gets a bad score, we look at the process rather than the writing. We look at previous drafts of the writer, their research notes, and the sources they used. We talk to the writers. The piece gets published, along with the report including the unflattering number, if the process is correct. It’s dishonest and makes the report meaningless to rewrite a good sentence to move a classifier.
Let us know if your check disagrees with our check after you submit it. Two revisions are available and can be used at any time over the next two weeks, so feel free to open a dispute. We will show you the drafts and notes, and will never use AI to humanize the draft to look more like an authentic human effort, because that is the only change that would augment the draft into being only machine-generated.
- Ask for the detector, the version used, and the threshold that yielded that result
- Request highlights at the sentence level, and review them for yourself
- Before reaching a conclusion, utilize a second tool that is different from an architectural standpoint
- Ask the writer for drafts, notes and sources, and read those too
Can AI Detectors Be Wrong FAQs
The output of every detector is only a probability of occurrence and part of a decision, which should be combined with other evidence. That is why Articled publishes reports and caveats simultaneously, because a check you cannot question is not a check.
The output of every detector is only a probability of occurrence and part of a decision, which should be combined with other evidence. That is why Articled publishes reports and caveats simultaneously, because a check you cannot question is not a check.
There is no universal answer that applies to your case as a number. The answer is dependent upon the tool, threshold, passage length, the genre, the author, and how recent the tool’s training is. Those saying one number applies to all texts are overselling their products.
Every text is scored differently, and the scoring changes rapidly with each vendor’s retraining or a new generation of models. We conduct twelve assessments rather than naming a single best option, and would be suspicious of any writing service that claimed a permanent winner.
There’s no published AI detector by Google and it says it rewards helpful, reliable content without considering the method of construction. If a third-party tool flags your page, it can’t be related to ranking. The problem is redundant and low quality content by whoever wrote it.
It is stronger evidence, not proof. A lot of detectors share the same architecture and training data, so errors correlate: the same non-native essay or the same boilerplate can trip a few tools for the same reason. Agreement doesn’t make something true; it just limits the doubt.
Better maintained and calibrated through retraining. This comes at a cost, but it’s worth the reliability you get. Reliability is not certainty, though. A tool costs money, but it gives you a probability from a statistical model, like a free tool would, and it can still misread heavily manipulated or edited dense prose.
Although the truthful position is often uncomfortable, it is important to advocate for it. These tools are great for preliminary filtering for large numbers of submissions, but should not be used as the exclusive tool for making a decision on an individual submission. Certainty about a single document should not be expected, as the selling point is unsupported by the mathematics.
Content a person actually wrote
$10 per 100 words, and the writer keeps all of it. Our 1% sits on top, 0.5% goes to trees, and no generated text appears anywhere in the process.