Blog

How Probator detects AI-written text, and how we avoid false accusations

The signals behind every AI-detection result, why we would rather say "inconclusive" than accuse a human writer, and where to find our accuracy figures.

An AI-detection result can affect a grade, a job application or a client relationship. So we built Probator around one rule: a wrong "AI-generated" is worse than an honest "inconclusive". This is how a check works.

Several independent signals

Guards against false positives

Formal, translated and non-native writing is the most often misjudged. Before we call a text AI-generated, three guards apply:

  1. Too short to judge. Under 80 words, we don't say "AI-generated" unless there is hard evidence.
  2. Two detectors must agree. If only one signal is confident, the result is "likely", never "AI-generated".
  3. Calibrated per language. Our model's threshold in each language is set so that at most 1 in 100 human documents in our validation data is flagged. If the model doesn't see AI writing, the result is "inconclusive".

When a guard changes a result, the report says so.

We publish our accuracy

Every model we put into use is tested on documents it never saw during training, and the results are on our accuracy page, by language. Our monthly reports show how results are distributed across all checks, anonymously.

What a result is not

It is a probability with reasons, not proof of authorship. If you assess other people's work, read the reasons, look at drafts and other evidence, and give the author the chance to explain. Our AI Policy sets out the rules.

Check your own text