Berk Bayri

Hallucination

When an AI model produces fluent, confident output that is false or unsupported by its sources. What matters is which errors are unacceptable in a given workflow and how they are detected.

A hallucination is output from a generative AI model that sounds plausible and is wrong, or cannot be traced to any supporting evidence. A language model predicts likely text; it does not check claims against reality unless the system around it is built to do so.

"Does AI hallucinate?" is one of the wrong questions about AI, because the answer is always yes, sometimes. The useful version is specific to a workflow:

  • Which errors are unacceptable here?
  • What evidence must the system provide for its claims?
  • How will unsupported output be detected?
  • What is the cost of a false positive versus a false negative?

Reducing and containing it

Grounding the model in retrieved sources with retrieval-augmented generation helps, but does not eliminate the problem, because the model can still misread or ignore what it retrieves. Add guardrails, require citations, route uncertain cases to a person, and track how often people have to repair a recurring error; that repair work shows up in supervision load. Good calibration helps a system know when it is not sure.

Read more in The wrong questions about AI right now.