← Back to blog

Two AIs gave me different answers: which one do I trust?

You ask ChatGPT something. Out of curiosity, you ask Gemini the same thing. They say different things. Sometimes it’s nuance; sometimes it’s the opposite.

The natural reaction is to assume one of them is broken, and go with whichever sounds better. Both of those are mistakes.

First: this is normal, not a malfunction

Two AIs disagreeing doesn’t mean one is faulty. They’re built differently, from different material, with different ideas about what makes a good answer. Ask something with one clear answer — how many days are in February — and they’ll agree. Ask something with nuance and there’s no reason they should.

In fact, if two different AIs agreed every time, you’d have a worse problem: it would mean asking two adds nothing, and that their agreement signals nothing.

Disagreement is what makes agreement worth something.

What the disagreement is telling you

When two answers split apart, it’s almost always for one of three reasons. Telling them apart is what turns confusion into something useful.

1. Your question doesn’t have a single answer. This is the most common case in real decisions. “Should I change jobs?” has no correct answer without knowing your savings, your tolerance for uncertainty, or what’s happening in your industry. Each AI filled those gaps its own way and landed somewhere different.

2. A fact is missing that only you have. The disagreement points at exactly which one. If one says “wait” and the other says “jump”, the difference usually rests on a specific assumption: how long you could go without income, whether you have another offer, whether the sector is growing. That’s the fact you need to supply.

3. One of them is making it up. Less common than people think, but it happens — especially with dates, figures, regulations or anything recent. And it comes unlabelled: an AI gets things wrong in exactly the same confident tone it uses when it’s right.

How to tell which one it is

You don’t need expertise. Three quick checks, in order:

Ask for the reasoning, not the answer. Ask each one what it’s basing this on. Seeing the reasoning side by side usually makes the fork obvious — and often reveals that they diverged over an assumption neither had mentioned.

Verify anything verifiable. If the answer contains a figure, a date, a law or a name, check it against a real source. Don’t argue with an AI about a fact: look it up.

Add what was missing and ask again. If you’ve identified the gap, put it in the question and re-run it past both. A lot of disagreements dissolve once the question stops having holes.

What not to do

Don’t go with the one you liked. This is the important trap. If you already wanted to quit, the answer saying “quit” will read as the sharper of the two. It isn’t — it’s the one that agrees with you. Using two AIs and then picking the one you already believed is worse than not asking, because you walk away with a false sense of backing.

Don’t split the difference. Two opposite recommendations don’t average. If one says “sign” and the other says “don’t”, the answer isn’t “sign half of it”.

Don’t keep asking until one agrees with you. Push hard enough and something will eventually tell you what you want to hear. That isn’t cross-checking: that’s recruiting an accomplice.

A real example

The demo on our homepage uses a case shaped exactly like this: someone on a permanent contract offered 15% more on a fixed-term one. One AI said take it; the other said stay. Both completely confident.

Neither was wrong. They were answering different questions, because each filled the unstated parts differently. We walk through it here, and the short version is that the disagreement was the useful answer: it flagged that the decision hinged on a fact only the person had.

When this happens a lot

Doing all of this by hand — two tabs, comparing, finding the fork, verifying the facts — is fine for one important decision a month. You won’t do it day to day, and you know it.

That’s precisely the work The Judge does for you. You write the question once, a jury of several AIs deliberates separately, and a judge — always the same one, same yardstick — delivers a verdict with its confidence level. When the jury doesn’t agree, it tells you instead of papering over it, because that’s exactly the part you need before deciding.

It won’t remove uncertainty where uncertainty is real. It removes the doubt about whether you’d have got a different answer somewhere else.

Start free