Verification

AI Detector Verification for Articled Content

Your drafts undergo twelve different tests for originality and potential AI assistance before leaving our office. Here, we explain what the tests measure, how to interpret the results, and what the tests fail to evaluate.

Order content How it works

How AI Detector Verification Supports Human-Written Content

Detectors utilize a statistical framework without actual witness accounts. They measure a passage one token at a time. Then, they assess how guessable each token was based on the preceding tokens. There are many valid options for each token, but AI-generated text typically selects a safe option repeatedly, resulting in less than human average surprise, or perplexity.

The second indicator is burstiness. This refers to variation in sentential length and complexity in texts. Imagine a writer produces a long, qualified sentence, then a short one. If AI models are used to write the text, they will naturally drift toward a simpler writing style. For example, if a tool claims to generate 92 percent AI in a text, it is not saying that 92 percent of the words in that text were generated using AI. It is stating how certain the classifier is that the text resembles the text that the tool was trained on.

This is a draft version of what we expect the finalized report to look like. We have the draft’s author on record, as we commissioned the writing and keep the writer’s notes. Here’s the electronic copy of the draft. It will be a time-saving artifact to add the draft to a file for a client, editor, or compliance reviewer. It would be helpful to avoid them performing a similar review and contacting you for clarification on answers you can’t give.

Perplexity and burstiness, briefly

Perplexity is a measure of surprise of a language model, while burstiness is a measure of text variability. If a classifier regards a text as machine-like, both measures are low and flat.

Which Tools We Use for AI Detector Verification

Twelve tools are found in the check, and this is not twelve different opinions for the same question. We are testing four different elements. Two of these elements relate to the pattern of the text. Are the patterns of the generated text visible? Is the wording similar to something that has already been published? We are also examining if the passage was quoted but was not indicated as a quoted passage. Lastly, we want to determine if there is any record of how the document was typed.

No single solution covers that entire area. Originality.ai and Copyleaks report AI scores alongside web-similarity checks. GPTZero, Winston AI, ZeroGPT, Sapling, Content at Scale, Writer.com and Crossplag are classifiers only, and have been trained on different data sets and configured with different metrics. Copyscape deals with duplicate content against a long-standing web index. Grammarly Authorship records ownership as opposed to probability and Turnitin is behind an institutional license, so a Turnitin result comes from the school or publisher that owns the license as opposed to from us.

The tools also change without notice. A file that looked clean in March can look differently in October, due to a change in the model that the tool is based on. Every report is dated, and the name of the tool is attached to the report. We treat an old report as a record of that day and not as an indefinite certificate.

  • The following tools are designed to identify whether text was produced by an LLM: GPTZero, Winston AI, ZeroGPT, Sapling, Content at Scale, Writer.com, and Crossplag.
  • Originality.ai and Copyleaks, both detection and plagiarism tools, provide two separate scores.
  • Copyscape checks for duplicated content through their public web index.
  • Grammarly Authorship is a tracking tool (typing and pasting) that establishes provenance.
  • Turnitin shows institutional similarity, as it is run by your client’s school or publisher that you are collaborating with.

Why AI Detector Verification Is Only One Part of Our Process

The only absolute is that no classifier can confirm a human wrote it. They can only say that the text does not appear to look similar to the writings included in their training data. For people with naturally even prose, it matters the most: technical writers, non-native English speakers, those with a tight house style, and those creating short factual text. It leaves a classifier almost no data to work with.

A flag, therefore, is the start of a conversation rather than a rewrite. An editor assesses the passage, compares it with the writer’s research file and previous drafts, and requests a meeting with the writer to explain how the section was constructed. If the work is genuinely theirs, we provide it as is, report the score, and do not waste time rephrasing until the score improves.

This is the entire policy. It is the same as the behavior that ends our working relationship with clients, and it is detrimental to the writing. The best evidence is located upstream: the author, a documented assignment, a trail of research, and no generative tools at any stage of drafting, editing, researching, or translating.

  • Proof of authorship, since a score describes style rather than who typed the words
  • A guarantee that another tool, or a newer version of the same model, will agree.
  • A trustworthy assessment of text less than a few hundred words long.
  • Anything search engines consult when they decide how to rank your page.

If a draft is flagged

Before delivery we tell you we found an issue. We provide the report that raised the issue and tell you what the editor found. We do not run the text through a humanizer or condense the text to improve a score.

AI Detector Verification FAQs

A human writer handles every part of this. Ask us on the contact page and a person will answer the same working day.

A human writer handles every part of this. Ask us on the contact page and a person will answer the same working day.

There is no individual score. Each vendor defines their own score and thresholds. We look for consensus. If eleven tools flag a draft as human and the other one does not, the one that does not will be read by an editor and not obeyed.

No, and it would be dishonest to say otherwise. Below a few hundred words most classifiers have too little signal to say anything meaningful, and many vendors state that themselves. On a 100-word order we run the checks that apply and note the length limit on the report.

No. Google does not run these tools. The reports are evidence for people: a client, an editor, a procurement team. They are not a signal that search engines read.

Typically we need a request to begin writing and public interfaces. We cannot access resources behind institutional licensing. Include any requests for reports with your order, and we will provide our frank opinion regarding report writing.

Send us the report. We will reprocess the file. We will then compare that result to the scans attached to the delivery, and explain the discrepancy. The discrepancy is usually the model is more current on the vendor’s side. Revisions are open for 14 days. We will also review a scan past the 14-day window.

Verification is the least expensive and least important part of the process we use. It determines if someone outside of the process assured the brief was read, the reading was completed, and a draft was written. It is possible to verify the process. The editorial workflow will answer any additional questions.

Content a person actually wrote

$10 per 100 words, and the writer keeps all of it. Our 1% sits on top, 0.5% goes to trees, and no generated text appears anywhere in the process.