Resources

Why Use Multiple AI Detectors?

A single score is a single model’s opinion after training with a single dataset and a single threshold. A panel transforms a verdict into a spread that you can read, argue about and audit.

Order content How it works

Multiple AI Detectors: The Basics

Detectors vary for valid reasons. They rely on different reference language models, have been trained on varying data comprising both human and machine writing, apply different algorithms to evaluate aggregate scores of sentences to assess the document, and owners of the detectors set thresholds according to various interpretations of what constitutes a worse error. By running the same sample through several of these, we are in fact creating several partly independent measurements as opposed to repetition.

“Partly” is the operative word. Many detectors are constructed with similar transformer architectures and trained on overlapping public data, leading them to share similar blind spots and errors. The sorts of texts one detector fails to classify correctly often trip up the others as well: texts by non-native English speakers, regulatory boilerplate, lists produced using templates, translated and short texts. In cases where there is agreement by the classifiers, doubt is reduced. However, this consensus should not be interpreted to mean that a blind spot shared by the classifiers has been corrected, since the detected blind spots are the result of the classifiers sharing the same training data.

There is hidden statistical cost with every additional tool. Each tool offers a new opportunity to cross a new threshold. Therefore, the more you add, the greater the odds that at least one of them will flag a document that is otherwise perfect. That is the reason why the panel must be considered a pattern rather than a maximum. One outlier in the presence of twelve others is to be expected. Nine out of twelve is a finding.

  • Whether several partly independent models converge on the same reading
  • The issue of which passages are flagged by multiple tools
  • The question of whether the output is a reliable signal or simply an artifact of an idiosyncratic tool
  • Whether the text has appeared elsewhere before, from the plagiarism tools

Why Multiple AI Detectors Matters

For anyone purchasing or commissioning writing, a single vendor PDF may as well be a single vendor’s opinion stated as a fact. A spread is a tool-based consensus with a record showing where tools disagreed and how. That makes the vendor’s strategy transparent. One report can be easily crafted. Providing a full spread, along with the one tool that showed an inconclusive or odd output, would mean the vendor has nothing to hide.

The panel asks different types of questions, and this is the part that most people overlook. For AI detection, the panel asks if the text looks like it was written by AI. With Plagiarism and duplicate-content detection, the panel asks if the text has been seen before, which is an index look-up rather than a probability evaluation. This is more reliable. With authorship provenance, the panel asks how the text was generated and entered in the document. Three different questions, three different types of evidence.

All of this comes with a cost. More tools bring more numbers, and subsequently, more discrepancies. This means more chances for a client to see a red bar and become concerned. We are willing to accept this trade, because the opposing option is a more neatly packaged report that provides less information. If a verification system only produces results that are contextually comfortable to the system, then it is not a verification system.

Every report goes in the email

We don’t choose the most appeasing. All results are included with your delivery, and if one of the tools contradicts the other eleven, you will see that contradiction, rather than hearing about it later from your own review.

How Articled Approaches Multiple AI Detectors

The panel is permanent and labeled, so you can check our work: Originality.ai, GPTZero, Turnitin, Copyleaks, Winston AI, ZeroGPT, Sapling, Content at Scale, Writer.com AI Detector, Crossplag, Grammarly Authorship, and Copyscape. Of these, ten of the tools measure whether text appears machine-generated. Grammarly Authorship tracks how the writing was entered into the document. Copyscape and the plagiarism detection features of Originality.ai check for non-original text.

An editor understands context. An algorithm does not. When a tool flags an issue, an editor reviews the surrounding text in context, compared to earlier drafts, and asks the writer about it. Often the flagged text is good writing that can’t be expressed any other way, like a statement of fact (definition) or a clarifying sentence (specification). There are sometimes real issues (Bugs) that the grammar checker is designed to catch.

We don’t adjust anything for the panel, and we can’t guarantee that every tool will return clean on every order. Clean tool returns would require what we ban. AI humanizers allow us to manipulate these stats without us changing the way we draft text. Using one categorically ends your engagement with Articled.

  • Guarantee that your detector will agree with our twelve
  • Prove authorship on its own, without drafts and notes behind it
  • Replace fact-checking, editing or a plagiarism review
  • Stay accurate forever as models and thresholds keep moving

Multiple AI Detectors FAQs

Instead of giving a single verdict, our software gives you a spread of opinion, and it is both more honest and more useful. Articled does all twelve reports for each of its drafts, and also submits those reports, even those which reflect poorly on them.

Instead of giving a single verdict, our software gives you a spread of opinion, and it is both more honest and more useful. Articled does all twelve reports for each of its drafts, and also submits those reports, even those which reflect poorly on them.

All twelve results are attached to this email, however, we also received results from tools that provided awkward numbers. We have included those results to you as well and you may ask us about anything else during the fourteen-day revision window.

An editor reviews the flagged portion and decides if it should remain in the final draft. The editor will reference any involved drafts and source notes. The writer and the editor often discuss the issue. Based on the analysis, the passage will remain in the final draft of the document, and the report will be included. We will not rewrite well phrased writing to boost a score.

No, and we can’t do that. More tools may mean more issues from ordinary human writing. What the panel buys you is the ability to have readable opinions and see how others disagree, instead of just having the final word given to you.

They are easier questions to address and to answer. Checking for duplicated content assesses your text with an index of published work. Unlike an AI evaluation of probabilities, that is a look-up with a definitive answer. Both can ruin an assignment and will both be evaluated.

Please include this on your order form, and we will do our best to accommodate. We have twelve tools standardized, so if your organization is using something else, let us know before the writer gets started. Although the writer will most likely know what to apply, it is more useful for the briefing stage to know what your standard is, so that the writer can be held against it.

The reason to run twelve is not Infinity War’s case of twelve scores compared to ten scores. That would be silly. It’s that twelve scores mean more work for them, you have more to work with to question their report, and there’s a higher chance of you finding a conflict worth exploring. A single score is probably the easiest thing in the world to forge, and therefore, the hardest to trust.

Content a person actually wrote

$10 per 100 words, and the writer keeps all of it. Our 1% sits on top, 0.5% goes to trees, and no generated text appears anywhere in the process.