You paste a paragraph, the meter stops at 78%, and nobody in the room can tell you what the 78 counts. Words? Sentences? Confidence? Guilt? That confusion is the normal state of AI detection in 2026: everyone quotes the number, almost nobody reads it correctly. This guide explains what sits behind the percentage so you can use it without letting it use you.
A percentage is not an authorship measurement
The number is an aggregate of probabilities, not a count of machine-written words. The tool splits your text into chunks, scores how predictable each chunk looks to a language model, and blends those scores into one headline figure. Nothing in that pipeline identifies who typed the keys.
Every vendor also sets its own bands. On one tool, 78% sits in a "likely mixed" band with a cautious explanation; on another, the same 78% renders as a red "likely AI" verdict. Same digits, opposite message. Comparing percentages across tools is like comparing temperatures without knowing which scale each thermometer uses.
The two signals behind the number
Nearly every detector blends two statistical signals. Burstiness is rhythm variation: human writing mixes short punches with long explanations, while steady, even sentences look machine-made. The second signal is word surprise (perplexity): predictable, common phrasing scores low, unusual word choices score high.
Both signals describe texture, not origin. Contracts, manuals, cover letters, and careful non-native writing all score low on surprise for honest reasons. A low surprise score says "this reads like average text", never "a machine wrote this". Keep that sentence in mind; the rest of this guide follows from it.
Why the same text scores differently everywhere
Four factors move the number more than authorship does. Length comes first: under roughly 200 words, detectors work with a handful of sentences and guess from coarse averages, so short samples swing wildly. Genre comes second: formulaic formats (five-paragraph essays, standard business mail, templated reports) share their shape with mass-generated text. Editing comes third: every polish pass smooths your rare-but-yours phrasing toward standard phrasing, which is exactly the statistic detectors flag. Calibration comes last: vendors tune thresholds differently, so disagreement between tools is expected, not a tiebreaker failure.
Test it yourself: run a three-page essay and one paragraph lifted from it. The long sample usually scores calmer than the paragraph, from the same author, in the same week. The number measured the sample size as much as the style.
How to read the number without getting it wrong
Set your threshold before testing, not after seeing the score. Decide in advance what "high" means for your context, require agreement between two tools before raising the question with anyone, and let a human reading make the final call. A detector is a reason to look closer, never a verdict.
And collect process evidence instead of chasing digits. Dated drafts, notes, sources, and revision history outweigh any percentage in a fair review. If you are the author, keep that trail as you write; if you are the reviewer, ask for it before concluding anything. Evidence of process beats a score in every dispute worth having.
If someone judged your text
Do not rewrite blindly to lower the number: chasing the score usually degrades clarity without moving the figure reliably, because each tool calibrates differently. Revise for readers instead. Vary sentence length on purpose, replace one generic passage with a specific example only you could give, and cut throat-clearing openers. These are the same edits that improve any text, flagged or not.
Vortixy helps with that half: paste the draft in the chat and ask for an honest review of rhythm, concreteness, and voice. You keep every decision, and the report explains each issue instead of hiding behind a percentage.
Try it in chat: review my draft for detector riskIf you judge other people's text
Never decide on a single score. Run the text through at least two tools, require agreement before speaking with the author, and always read the work first yourself. Write the rule down so it applies to everyone: one score triggers a closer look, two agreeing scores plus a human reading trigger a conversation. Teams that codify this have fewer disputes and fairer outcomes.
Frequently asked questions
Is there a percentage that proves AI authorship?
No. Scales differ per tool, and both underlying signals shift with length, genre, and editing. Treat any fixed cutoff as a policy choice your team made, not a fact the tool discovered.
Why did my fully human text score 80%?
The usual causes are a short sample, a formal or technical genre, heavy editing that evened the rhythm, or careful phrasing that lowered surprise. Lengthen the sample and evaluate the longer evidence before concluding anything.
Should I rewrite until the score drops?
No. Score-chasing degrades writing without lowering the number reliably. Revise for readers (rhythm, concreteness, voice) and present process evidence if anyone questions the work.
Do longer samples score more fairly?
Usually yes. Short samples starve the statistics, so the tool guesses from coarse averages; long samples let real variation appear and scores stabilize. If one paragraph flags and three pages from the same author do not, trust the long evidence.