Detecting AI / Evidence desk

Why AI Detectors Give Different Results

Why AI detector scores disagree, what the numbers actually mean, and how to compare results without cherry-picking a verdict.

Oct 1, 20253 min readEvidence desk

Run the same document through two AI detectors and the results may differ. That does not automatically mean one result is fake or one system is broken. It usually means the systems are measuring related signals with different models, thresholds, and definitions.

The disagreement matters because a score can look more precise than it is. Before comparing two numbers, check what each number represents.

The scores may answer different questions

One product may report a document-level likelihood. Another may estimate the share of sentences that match its AI pattern. A third may place the text into broad labels such as low, mixed, or high signal. A score of 72 in one system is not necessarily comparable to 72 in another.

Read the label, help text, and methodology before treating a number as a probability. If the product does not explain what its score means, the number is difficult to use responsibly.

Five reasons results change

1. Different training data

Detectors learn from collections of human and generated text. Those collections may cover different models, languages, genres, and dates. A detector trained heavily on essays may react differently to marketing copy or technical documentation.

2. Different model features

Systems can look at word choice, sentence variation, predictability, repetition, transitions, or relationships across a longer passage. They may weigh the same feature differently. Two models can therefore observe the same text and reach different conclusions.

3. Different decision thresholds

A detector must decide how much evidence is enough for a high-signal label. A strict threshold may reduce false positives but miss more generated text. A sensitive threshold may catch more generated text while flagging more human text. There is no threshold that removes both kinds of error.

4. Different text preparation

Formatting, headings, citations, bullet lists, quotations, and reference sections can affect an analysis. Some systems remove this material before scoring. Others process it as submitted. Even a minor edit can change a result when the text sits close to a decision boundary.

5. New models and editing methods

Generated text changes as language models change. Human editing, translation, and rewriting also alter its statistical pattern. Research has shown that detection becomes less reliable under practical transformations and adversarial rewriting. A model's past benchmark does not guarantee the same performance on every new text.

How to compare two results

  1. Use the exact same text. Keep formatting and excluded sections consistent.
  2. Record the date and settings. Detector models can change, so a result is a snapshot.
  3. Compare labels, not just numbers. Confirm whether each score describes the whole document, selected sentences, or model confidence.
  4. Inspect the passages. Sentence-level evidence is more useful than a single unexplained total.
  5. Do not shop for a preferred verdict. Repeating checks until one score supports a desired conclusion is not a sound review method.

What disagreement tells you

Close agreement between independent systems can strengthen a reason to investigate, but it still does not prove authorship. Strong disagreement is a warning that the text may sit near a boundary, fall outside a model's strengths, or be affected by preprocessing. In either case, the next step is human review.

Use the methodology page to understand how Detecting AI presents its signals. Our result is designed to point a reviewer toward relevant passages. It is not a substitute for drafts, sources, policy context, or a conversation with the writer.

Sources

Detecting AI Evidence Desk

We publish practical guidance for careful authorship review. Detection scores are signals to investigate, not proof of misconduct or authorship on their own.