← Back to blog

An AI that tells you when it isn't sure

An AI answers just as well when it knows as when it’s making it up. Same prose, same poise, same tidy structure. No wobble in the voice.

That’s the real problem. Not that it fails — they all do, and so do people — but that it fails without changing its tone. A human expert says “you’d want to check that” or “not my field”. An AI, by default, doesn’t.

Why it doesn’t hesitate

It isn’t malice or sloppiness. These systems are trained to produce the most plausible continuation of a text, and plausible text is confident text. Hedging isn’t rewarded.

Ask it to commit and it commits. Ask about something it knows nothing about and it will build something with the shape of a good answer. That’s what it’s good at.

You can ask how sure it is, and it will tell you. But that number has a problem: it’s being generated by the same machinery that generates everything else. An AI grading itself tends to give itself good marks — not because it’s lying, but because it has no way of knowing what it doesn’t know.

Where a useful doubt can come from

If the model can’t reliably measure its own certainty, the doubt has to come from somewhere else. And there is somewhere: ask several times, of different systems, and see whether they agree.

It’s what you already do when something matters. You get three quotes. You get a second medical opinion. Not because the first one lies, but because agreement between independent sources says something no single source can say about itself.

When three AIs from different providers say the same thing, there’s solid ground underneath. When they say opposite things, you’ve learned something you’d never have learned from one: that your question has no automatic answer.

We’ve seen it in real cases. We asked several AIs about a rent increase and some said the landlord could, another said he couldn’t. None was lying: the answer depended on facts that weren’t in the question, and each quietly assumed a different scenario. The disagreement was the information.

What a meaningful confidence level looks like

Here’s the difference between a decorative number and a useful one:

Useless: a percentage produced by the same model that wrote the answer. It’s asking someone to mark their own exam.

Useful: a number that reflects two measurable things — how much independent answers agree, and how much of what they say is verifiable rather than opinion.

That’s how ours works. A fixed judge reads what each juror said and grades by the same rubric every time: if they agree and it’s verifiable, high confidence; if they differ on something material or the question is ambiguous, low confidence and a reason. A 68 isn’t decoration — it means there’s a real disagreement and you should look at it before deciding.

Which is also why we refuse to round it up. A verdict that always said 90 would be easier to sell and worth nothing.

How to spot it yourself, with no special tools

Even if you never use anything like this, there are signs an answer is more fragile than it sounds:

  • No caveats at all. An answer to a complex question without a single “it depends on” is suspicious. Reality has conditions.
  • Specific facts with no source. Figures, percentages, legal articles, names of standards. It’s where invention is most common and most convincing.
  • It asks you nothing. If your question needed a fact you didn’t give and it answered anyway, it assumed one.
  • It changes if you ask again. Ask the same thing in a fresh conversation. If the answer shifts substantially, the first one wasn’t a certainty: it was one of the options.

That last trick is the cheapest and the one more people should use. It costs a minute.

What to demand from an AI

Not that it’s always right. Nobody is.

What you should demand is that when it doesn’t know, it shows. That the difference between “this is well established” and “this is my best guess” reaches you, instead of staying inside.

It’s the only thing that turns an answer into something you can decide with.