Asking multiple AIs at once: how to do it properly
The short version: different models, the same question word for word, none of them knowing what the others said, and compare the reasoning rather than the conclusion. Miss any of those four and you have three answers but no comparison.
Asking several AIs at once has become a sensible reflex: if one can be confidently wrong, let three answer and see what comes out. The instinct is right. The execution is almost always broken, and in a way you don’t notice — because the result looks like a comparison.
The four conditions
1. Different models, not different tabs. Two windows of the same model aren’t two opinions: it’s the same training answering twice. They’ll agree, and that agreement tells you nothing. If you’re going to spend the effort, make them different houses.
2. The same question, literally. Retyping it from memory in the second tab seems harmless and ruins the whole thing: you’re no longer comparing two answers to one question, you’re comparing two answers to two questions. Copy and paste, every time.
3. Keep them uncontaminated. The moment you tell the second one “another AI said X”, it stops being an independent opinion and becomes a reaction. These systems tend to accommodate whatever is put in front of them: they’ll correct a detail and accept the frame. You’ve lost exactly what you came for.
4. Compare the reasoning, not the headline. Two identical conclusions reached for opposite reasons are an alarm, not an agreement. And two different conclusions can hide the same reasoning with a different emphasis. The headline carries the least information of anything on the page.
What you’re actually looking for
Not a vote count. The exact point where they part ways, which is nearly always one of three things: a fact (one of them is objectively wrong and you can check it in two minutes), an assumption (each filled in something you didn’t say, in its own way), or a value judgement (they weigh stability and upside differently, and that one is yours).
We go into it in how to compare two AI answers, which is this post’s sibling.
How many, and when to stop
Two is enough to detect there’s a problem. A third breaks the tie when the first two disagree on something that costs money or is hard to undo. Beyond that, more answers don’t add information: they add reading and a false sense of rigour.
And if all three agree, raise your confidence but don’t call it a guarantee: three models trained on the same wrong page will agree quite happily. What comparison removes is the lone mistake, which is the common case.
Why almost nobody sustains it
Doing this properly — different models, identical question, uncontaminated, reading each one’s full reasoning — is about twenty minutes a question. You do it the first time, do it roughly the second, and abandon it by the third, which is exactly when the decision matters and you’re in a hurry.
That’s why this exists: convening several AIs, keeping them independent and having a fixed judge compare the substance is precisely what you’d do by hand, done the same way every time and without depending on your energy that day. Plus something you can’t get by hand: a confidence level measured with the same rubric on every question.
Frequently asked questions
Can I use ChatGPT and an older version of it? No. They share training and biases: it’s nearly the same opinion twice.
What if one is much better written than the others? Writing quality isn’t reasoning quality. It’s the most common mistake when comparing by hand.
Do I have to do this for everything? No. For a recipe or a first draft, one is plenty. Save it for what costs money, time, or can’t be undone.
What if they disagree and I can’t tell who’s right? That’s already a result: it means your question depends on something you didn’t say, or is genuinely uncertain. It’s what stops you deciding with a confidence that was never there.