AI detectors rarely explain themselves. They hand you a percentage and leave you guessing what was measured. Under that number there are usually two statistical signals with odd names: perplexity and burstiness. Understanding them changes how you read every detector score — and why the same text can look human in one tool and machine-made in another.
Predictability: how surprising is each word?
In technical writing this measure is called perplexity: it tells you how surprised a language model is by each word. A sentence built from common, expected word choices has low perplexity: the model saw it coming. A sentence with unusual words, rare combinations, or sharp turns has high perplexity: the model did not expect it.
AI-generated text tends toward low perplexity because models pick likely words by design. Human writing spikes and dips — a plain sentence, then a strange one, then jargon only an insider would use. Detectors treat sustained low perplexity as a machine fingerprint.
The catch: predictable is not the same as artificial. Legal boilerplate, technical documentation, and beginner essays all read as low-perplexity because the genre demands standard phrasing. A detector cannot tell the difference between "written by a model" and "written in a formula".
Burstiness: where is the rhythm?
Burstiness measures variation — sentence length, complexity, and structure across a passage. Humans burst: a three-word punch, then a long winding explanation, then a fragment. For effect. Models drift toward the middle: medium sentences, medium complexity, steady rhythm, paragraph after paragraph.
This is the signal most people feel before they can name it. Read a suspect text aloud. If every sentence lands with the same weight and length, burstiness is low — and most detectors will lean toward machine-made.
But uniformity has human causes too. Tight editing flattens rhythm. Style guides enforce it. Non-native writers aiming for correctness produce careful, even sentences. Each of these reads as "low burstiness" without any machine involved.
Why both signals shift
Neither signal is stable across contexts, which is why percentages swing:
- Length. Short samples give the statistics almost nothing to work with. A single paragraph can score anywhere; the same text embedded in three pages settles down.
- Genre. Poetry and chat messages burst wildly; contracts and manuals do not. A detector tuned on essays misreads both ends.
- Editing. Every revision pass — human or machine — smooths the text toward the average. Heavily polished human writing converges on the same statistics as generated text.
- Topic familiarity. Writing about a well-covered topic pulls in standard phrasing (low perplexity). Writing from fresh experience breaks patterns (high perplexity). Same author, different scores.
Reading the two signals together
Three rules keep perplexity and burstiness useful instead of misleading:
- Compare scores only within the same tool. Each detector calibrates its scales differently, so a number from one is not evidence in another.
- Retest before acting. Run the same text twice, and run a longer sample when the first one is short. Scores that swing between runs were never verdicts.
- Set the threshold before you test. Decide what happens above and below your line — manual review on both sides — instead of negotiating with the number afterward.
Vortixy and the review-first approach
The productive response to a strange score is not a better detector. It is a better draft: checkable facts, one idea per paragraph, concrete phrasing instead of filler, and a rhythm that varies on purpose. Vortixy reviews the draft for clarity, structure, voice, and factual consistency, explains each issue, and returns a revision you accept or adjust. You keep the decision; the report shows the work.
Try it in chat: review my draft for flat rhythm and vague phrasingFrequently asked questions
What is burstiness in AI detection?
Burstiness is the variation in sentence length and complexity across a text. High burstiness — short punches mixed with long explanations — reads human. Sustained evenness reads machine-made to most detectors, though editing, style guides, and careful non-native writing also flatten rhythm.
What is perplexity in AI detection?
In detection, perplexity is how predictable each word is to a language model. Low perplexity means the text uses expected, common phrasing; high perplexity means surprising word choices. Generated text skews low because models prefer likely words — but so do contracts, manuals, and formulaic human writing.
What AI detection percentage is bad?
There is no universal bad percentage because scales differ per tool and both underlying signals shift with length, genre, and editing. Set your own threshold before testing, require agreement between two tools, and confirm with a human reading before any consequential decision.
Why does my human writing flag as AI?
The usual causes are short samples, formal or technical genres, heavy editing that evened the rhythm, and careful phrasing that lowered perplexity. Lengthen the sample, check whether the genre itself is formulaic, and review the text for quality — publishable and machine-made are different questions.