How AI text detection works, and when it is wrong
In shortAI text detectors estimate how likely a passage was written by a language model. ProofVerity does this sentence by sentence, reports a calibrated probability with its error rate, and shows the signals behind every flag. No detector is proof on its own; the evidence is what you act on.
What an AI detector measures
Language models write with unusually predictable word choices and an even sentence rhythm, and they lean on a recognisable set of phrases. A detector measures these properties against a baseline of human writing for the same kind of document.
- Perplexity: how surprising each word is given the words before it. Machine prose is rarely surprising.
- Burstiness: how much sentence length and structure vary. People vary; models flatten.
- Marker phrases: constructions that models over-produce ("it is important to note", "delve into", "pivotal role").
- Style breaks: a passage whose statistics differ from the rest of the same author’s document.
The output is a probability that the span was machine-written. It is not a fact about the author, and it is never reported by ProofVerity as a single score for a whole document.
Why ProofVerity works at sentence level
Most real documents are mixed: a human draft with a machine-polished paragraph, or a machine draft with human edits. A whole-document score hides that mixture and cannot be discussed. A sentence-level flag names the exact passage, which lets an author, a student or a source respond to something concrete.
Sentence-level detection also lets ProofVerity compare a passage with the author’s own baseline in the same document. A sudden change in rhythm and vocabulary is evidence in itself, and it is shown as such.
What a calibrated probability means
A 94 means 94: among all spans ProofVerity marks at 94%, about 94 in 100 were machine-written in evaluation. Calibration is done separately for each document type, because a lab report, a news brief and an admissions essay have different baselines, and a detector tuned on one will misfire on the others.
Every report states the false-positive rate for its document type next to the probability. The rate is the share of human-written spans that the detector would wrongly mark at that threshold. It is published so that a reader can decide how much weight a flag deserves.
When detectors are wrong
All AI detectors make mistakes, in both directions. The situations below are where false positives concentrate, and what ProofVerity does about them.
- Non-native and formulaic writing reads "flat" to a naive detector. ProofVerity calibrates per genre and discounts short spans, and it never reports a whole-document verdict.
- Templates and boilerplate (methods sections, legal wording, lab protocols) are predictable by design. Style-break detection against the author’s own text separates a template from a machine paragraph.
- Heavily edited machine text keeps some rhythm signals but loses marker phrases; confidence drops and the report says so.
- Paraphrasing tools change words but rarely restore human variation, and they cannot make an invented reference exist. The citation check catches what the text check misses.
- Translated text carries the statistics of the translation engine. Reports note when a document appears to be a translation.
How to read a result
- Look at the span, not the document. Which sentences are marked, and how many.
- Read the evidence lines: perplexity against the human median, burstiness, marker phrases, style break.
- Check the false-positive rate for this document type.
- Look at the citations. Invented or misquoted references beside a flagged paragraph change the picture; verified ones do too.
- Ask before you accuse. Drafts, notes and a short conversation settle most cases.
- Record the decision: mark reviewed or dismiss, add a note, export the annotated PDF with its evidence appendix.
What ProofVerity does not do
- It does not issue verdicts or "AI scores" for whole documents.
- It does not compare a text with other students’ work; it is not a plagiarism checker.
- It does not train on your documents, and it deletes them after 30 days.
- It does not hide its reasoning. Every flag shows the signals that produced it.
Questions
- Can AI detectors be beaten by paraphrasing tools?
- Partly. Paraphrasers change surface words but usually keep the flat rhythm, and they cannot repair invented references. ProofVerity weights rhythm and structure and runs the citation check alongside. No detector is airtight, which is why results are evidence rather than verdicts.
- Does ProofVerity detect text from ChatGPT, Claude, Gemini and other models?
- It detects machine-written prose regardless of the model, because it measures properties of the text rather than one model’s fingerprint. New models are added to the calibration sets as they appear.
- Is a high probability enough to fail a student or reject an article?
- No. Use it to start a conversation, together with the citation check, drafts and a short follow-up. ProofVerity is designed so that the person whose name is on the decision makes it.
- Where are the error rates published?
- In every report’s evidence appendix, next to each probability, and in the Method section of the site. They are re-evaluated whenever the calibration sets are updated.