Berk Bayri

The AI should answer every question

A common misconception about AI models and evaluation, tested against the evidence.

The myth
A useful AI system should try to answer every question instead of refusing, asking for clarification or saying it does not know.
The reality
In many real decisions, a calibrated refusal is better than a plausible guess. A system that knows when to abstain can be more useful than one with a slightly higher raw answer rate.

Explanation and evidence

The argument

A system that always answers can look more capable than a system that sometimes stops. That comparison is often backwards.

OpenAI's research on makes the incentive problem unusually concrete. When models are scored mainly on whether they produce the correct answer, guessing can outperform abstaining even when the guess is wrong most of the time. In one SimpleQA comparison cited by OpenAI, an older model had slightly higher raw accuracy but a dramatically higher error rate because it almost never abstained.

That is not merely a curiosity. In a business process, the cost of a confident wrong answer can be much larger than the cost of asking for clarification, retrieving more evidence or escalating to a person.

A system that answers less can make better decisions if it knows which answers are too expensive to guess.

Why people believe it

Chat interfaces teach us to treat silence or refusal as failure. Users ask; software responds.

But decision systems have another option: do not decide yet.

What the evidence says

OpenAI and Anthropic evaluations show a real tradeoff between coverage and confident error. Some models answer more questions; others refuse more often and hallucinate less. Neither extreme is automatically best. The right operating point depends on the consequence of being wrong.

For a low-stakes brainstorm, guessing may be acceptable. For a payment, medical recommendation, compliance decision or irreversible agent action, uncertainty should change what the system is allowed to do.

The better question

Do not ask only, "How often does the AI answer correctly?"

Ask: When should it answer, when should it retrieve, when should it ask, and when should it ?

That is not a limitation to hide. It is part of the product.

Sources

Why language models hallucinate

OpenAI · 2025-09-05

Findings from a pilot Anthropic–OpenAI alignment evaluation exercise

OpenAI · 2025-08-27

Related reading

Related misconceptions