A detector score is an estimate produced by a particular system. It should not be treated as proof of authorship or a measure of writing quality.
A detector labels a paragraph “82% AI.” The number looks precise, but its meaning depends on the tool. It might describe the share of qualifying text flagged by the system, a classification score or another vendor-defined measure.
Before reacting to the percentage, find out what it actually represents. It is not automatically an 82 percent probability that a particular person used AI.
Different scores answer different questions
Turnitin, for example, describes its overall percentage in terms of qualifying text that its system identifies as potentially AI-generated or modified by certain AI tools. That is different from a plagiarism similarity score.
Other detectors may define their outputs differently. Read the explanation provided by the specific product. Do not assume two percentages can be compared as though they came from the same measurement.
A score with no clear definition is difficult to use responsibly, no matter how polished the interface looks.
Errors can go in both directions
A false positive occurs when human-written material is flagged as AI-generated. A false negative occurs when AI-generated material is missed.
Turnitin warns that its model can misidentify writing and should not be the sole basis for adverse action against a student. NIST's evaluation work also shows that performance depends on the generator and detector being tested. That does not mean every detector is useless; it means context and evaluation matter.
A result obtained on one collection of writing cannot establish perfect accuracy for every language, subject and writing style.
Short passages create a narrower view
A short paragraph gives a tool less material to assess than a long document. Some products set minimum input requirements or limit which kinds of text qualify.
Follow those requirements. A list of product specifications, a table or a few conventional sentences may not fit a detector's intended use. The absence of a warning on the screen does not tell you that the test is meaningful.
Treat unexplained changes between repeated scans as a reason to investigate the tool, not as evidence that authorship itself has changed.
Better evidence comes from the writing process
If authorship matters, consider drafts, notes, source records and revision history. Ask the writer to explain their reasoning and the decisions behind the work.
These records need context too. A saved draft does not answer every question, and missing notes do not prove misconduct. The aim is a fair review using several relevant pieces of evidence.
For publishers, the more useful editorial questions are whether the claims are accurate, sources are represented fairly and the article helps the intended reader.
Do not edit only to move a meter
Awkward wording, deliberate mistakes and unnecessary sentence changes can make a useful article worse. A lower detector score does not repair an unsupported claim or add a missing explanation.
Revise for clarity, specificity and evidence. Add an example that answers a real reader question. Remove a paragraph that repeats the introduction. Check any statistic you cannot trace.
A detector may contribute one signal to a review. It cannot replace a clear editorial standard or provide a universal guarantee about who wrote the text.
Sources & further reading
Original explainers and practical examples, with technical background from the sources below. Source links reviewed 2026-10-03.
More context, fewer assumptions. About our editorial approach.

