Citations mean the AI answer has been verified: a common AI evaluation misconception
RealityA citation can be real and still fail to support the sentence beside it. Links make verification possible; they do not perform the verification for you.
A common misconception about AI models and evaluation, tested against the evidence.
An AI detector does not observe how a document was created. It observes the document after the fact and estimates whether its statistical features resemble examples it has learned to classify as AI-generated.
That is a very different claim from authorship.
OpenAI withdrew its own text classifier after describing its accuracy as too low for reliable use. Stanford researchers found another failure mode: several detectors disproportionately labeled essays by non-native English writers as AI-generated. More recent evaluations continue to find that detector performance changes with model generation, editing, paraphrasing, language and the mixture of human and AI text.
A detector can classify a text. It cannot witness who wrote it.
The interface usually produces a percentage, label or colored score. Numbers feel forensic.
The word "detector" reinforces the impression that there is a stable fingerprint being found, like malware or a chemical residue. But modern generated text does not carry one universal visible signature, and humans can edit AI output while AI can imitate many human styles.
Detector performance can be useful in controlled settings, especially when the model family, text length and domain resemble the detector's evaluation data. That does not turn the result into proof for an individual document.
False positives matter most when the consequence is disciplinary, reputational or professional. A tool can be statistically useful across a dataset and still be unsafe as a verdict on one person.
Instead of asking, "Did AI write this?", ask:
What evidence do we have about the creation process?
Draft history, version history, source notes, document metadata and a conversation with the author can establish provenance more directly than a detector score alone.
RealityA citation can be real and still fail to support the sentence beside it. Links make verification possible; they do not perform the verification for you.
RealityA benchmark measures performance under its own task distribution and rules. Production adds messy inputs, tools, latency, repeated runs, changing workflows, failures and consequences the benchmark may never test.