AI detectors promise a simple answer: paste text, get a percentage, know if it was written by a machine. In practice the percentage changes between tools, between uploads, and even when the same text is tested twice. This is not a flaw in one product. It is how the underlying technology works.
Most detectors are trained on patterns from known AI outputs and human writing. They compare the target text against those patterns and return a probability. That score depends on the training data the detector used, the style of the source text, and how the AI that produced it was configured.
The same paragraph can score high in one tool and low in another. Short texts score less reliably than long ones. Heavily edited drafts drift away from the patterns the detector learned. None of this means the tools are useless. It means a single number is not a verdict.
Detectors measure statistical similarity to their training data, not intent and not authorship. A text written by a person who happened to use formulaic phrasing can trigger a high score. A heavily edited AI draft can score as human.
If you must use a detector, treat it as one weak signal among many. Never publish, reject, or accuse based on the score alone. The cost of a false positive is a real person being told their work looks machine-made.
The more useful question is not "did a machine write this?" but "is this text accurate, clear, and ready to publish?" That is a review problem, not a detection problem.
Vortixy approaches it that way: it reviews the draft for clarity, structure, voice, and factual consistency, and returns a revision with the issues explained. You keep editorial control while the report shows where the text is weak.
A practical checklist works alongside any tool:
There are legitimate cases for detection: auditing content pipelines, checking academic submissions against institutional policy, or investigating whether a document was machine-generated. In those cases the tool should support a human process with clear limitations, not replace judgment.
Set a threshold before testing, run the same text through at least two tools, and treat disagreements as a reason to review the text manually rather than to pick the score you prefer.
A detector percentage is a similarity score, not a share of the text that is machine-made. A "62% AI" label does not mean 62% of the sentences came from a model. It means the text resembles the detector's AI training examples strongly enough to score 0.62 on its internal scale.
Three rules make percentages useful instead of misleading:
No public benchmark settles which detector is "most accurate", because accuracy depends on the test set: language, length, genre, how much the draft was edited, and which model family produced it. A detector tuned on English essays can stumble on short product copy. One tuned on one model family can miss text from another.
That is why honest comparisons report their method — sample size, text lengths, edit levels — instead of a single winner. If a comparison hides its test set, treat its ranking as marketing. Build your own mini-benchmark instead: ten texts you actually handle, each run through two tools, with disagreements resolved by human review. Ten real samples beat a hundred synthetic ones for your decision.
AI detection is probabilistic, not certainty. The score you see today may differ tomorrow, and it says little about whether the text is good. If your goal is publishable, trustworthy content, spend the review effort on the text itself. A clear revision process with human judgment beats a percentage every time.
They are moderately useful signals, not verdicts. The same text can score high in one tool and low in another, short texts are less reliable than long ones, and heavy editing moves scores in either direction. Use detectors as one weak signal inside a review process, never as the decision.
There is no universal bad percentage because scales differ per tool. The workable approach is to set your own threshold before testing — for example, review manually everything above the level you chose — and to require agreement between two tools plus a human reading before any consequential decision.
Each detector trained on different data, weighs different patterns, and calibrates its scale differently. Style, length, edits, and the source model's configuration all shift the score. Disagreement between tools is expected and is itself a signal to review the text rather than to trust either number.
Academic workflows that use systems such as Turnitin should follow the institution's policy and treat any AI score as one input among several. Run the text through the approved process, document the threshold used, and resolve flagged cases with instructor review of the work itself — drafts, sources, and revision history — instead of the percentage alone.