An AI detector score is not a chain of custody. It is a statistical signal produced by a tool with an error rate that varies by text length, genre, language background, editing history, and model version.
This AI detection evaluation guide provides a fairer workflow for schools, publishers, and product teams. It builds on the Declaration of Independence false-positive example and the broader AI chatbot detection methods article.
Start with the actual question
“Was AI used?” is often too vague to govern. A useful review separates at least four questions:
- Was an AI system used at any point?
- Was its use allowed for this assignment or workflow?
- Did the person disclose or document that use as required?
- Is the submitted work accurate, original, and attributable to the person?
A detector score cannot answer all four. It cannot identify the exact tool, prove intent, or explain whether an allowed grammar correction became an undisclosed ghostwritten submission.
Evidence hierarchy
Prefer evidence that describes the process rather than a number inferred from the final text.
| Evidence | What it helps establish | Main limitation |
|---|---|---|
| Drafts and version history | How the work changed over time | Some tools do not retain history |
| Notes, sources, and citations | Whether the author understands and supports the claims | Notes can also be generated or incomplete |
| Conversation or tool logs | What an AI system was asked to do | Logs may be unavailable or privacy-sensitive |
| Author interview or walkthrough | Whether the person can explain decisions and revise the work | Human judgment is still fallible |
| Detector score | A reason to ask more questions | False positives and tool drift |
The correct outcome is usually a confidence assessment with documented uncertainty, not a binary label produced by one scan.
Why detector scores fail
Detection tools look for patterns associated with generated text. Those patterns also appear in human writing that is formal, highly edited, formulaic, translated, or produced by someone writing in a second language.
Short passages are especially unstable because there is less evidence to distinguish style from chance. A score can also change when the same text is lightly edited, reformatted, or sent to a different detector version.
For that reason, never compare scores from different tools as if they were calibrated probabilities. Record the tool name, version, date, input length, language, and whether the text was edited before scanning.
A four-stage review workflow
1. Preserve context
Save the submitted text, the detector report, and the conditions under which it was produced. Do not paste sensitive work into unapproved third-party services just to obtain another score.
2. Request provenance
Ask for drafts, source material, outline notes, citations, and a brief explanation of the editing process. In a workplace, repository history, design files, and review comments can provide the same kind of evidence.
3. Conduct a human review
The reviewer should check factual accuracy, source quality, consistency with prior work, and the author's ability to explain the result. The review should not be a disguised interrogation based on writing style.
4. Offer an appeal path
Document what evidence changes the decision, who can review an appeal, how long records are retained, and how privacy is protected. An opaque automated rejection is not a defensible policy.
Designing a responsible policy
A responsible policy should say:
- which AI uses are allowed, restricted, or prohibited;
- what disclosure is required;
- what process evidence people should retain;
- that detector scores are never conclusive on their own;
- who reviews a flagged case;
- how the person can respond or appeal;
- how sensitive drafts and logs are stored and deleted.
The policy should also be tested on known human writing, translated writing, and edited writing before it is used in a high-stakes decision. If the false-positive rate is not acceptable, the detector should not be used as a gate.
Bottom line
AI detection works best as a prompt for a fair conversation about process, quality, and policy. It becomes harmful when a probabilistic score is treated as proof. Preserve provenance, review the work, explain uncertainty, and keep a meaningful appeal path.