← Back to blog

ChatGPT can be confidently wrong: how to catch it before you decide

Ask ChatGPT, Gemini, or any AI something it doesn’t actually know for sure, and it will almost never tell you that. It will give you an answer. In the exact same confident tone it would use to tell you what 2+2 is.

That’s the underlying problem, not an occasional glitch: an AI doesn’t know what it doesn’t know. There’s no internal alarm that goes off when it’s improvising. And when the question actually matters — a contract, a job change, a big purchase — that unqualified confidence can get expensive.

Why an AI never says “I don’t know”

An AI generates the most probable answer based on what it’s learned, not the truest one. Most of the time those two things line up. But when they don’t, there’s no visible signal telling them apart: the same confident tone covers what it knows cold and what it’s making up.

That’s why “just ask the AI” isn’t a bad idea — it’s an incomplete one. It’s missing the follow-up question: and if it’s wrong, how would I even know?

Three signs you should double-check

You don’t need to be an expert to spot them. With practice, these patterns show up again and again:

  • It gives you a long answer to something that should be simple. Sometimes length is padding for real certainty it doesn’t have.
  • It never hedges. A genuinely expert human answer almost always includes an “it depends.” A genuinely unsure AI almost never does.
  • It flips its answer if you push back. Ask the same thing a different way, or say “are you sure?” If it reverses itself with the exact same confidence as the first answer, neither version was actually calibrated.

The trick that actually works: a second opinion

The most reliable signal isn’t inside a single answer — it’s in comparing two. Ask the same question to two different AIs (ChatGPT and Gemini, say) and see what happens:

  • If they agree, you have a real reason to trust it.
  • If they disagree, you just found the exact point where you should think twice — before deciding, not after.

It’s the same principle we already use outside of AI: a second medical opinion, a second quote, a second pair of eyes on a contract. With a single AI, you lose that safety net without noticing, because its answer sounds complete even when it isn’t.

When this actually matters

You don’t need to cross-check everything. For “what’s the capital of Portugal,” it doesn’t matter. But before a decision that costs money, time, or is hard to undo — negotiating a salary, choosing between two offers, deciding whether a big purchase is worth it — that extra minute of comparing is cheap next to the cost of being confidently wrong.

That’s exactly the idea behind The Judge: instead of convening several AIs yourself and comparing their answers on your own, it does it for you — gathers the best ones, makes them deliberate, and gives you a verdict with its confidence level, warning you when even they can’t agree.

Get early access

Prototype · simulated data