What a jury of AIs is (and why it isn't a multi-model chat)
When you explain that several AIs answer the same question, the usual reaction is: “ah, like those sites that show you four answers side by side”.
They’re alike for the first half and diverge exactly where the value is. The difference fits in one question: who compares the answers?
A comparator gives you work. A jury gives you an answer.
A multi-model comparator does something useful: it sends your question to several AIs and shows you what each one said, side by side.
And there it leaves you. With four long texts, confidently written, saying similar things with different nuances, and the job of comparing them is yours. Which is precisely the hard job.
What happens in practice is predictable: you read the first one properly, skim the second, and keep the one that’s best written or the one that agrees with what you already thought. You paid for four answers to end up using one, chosen on criteria that aren’t quality.
What a jury adds
Three things, and all three are what you can’t sustainably do by hand:
Independence. Each AI answers without seeing what the others said. It sounds obvious and it’s the first thing that breaks when you do it yourself: the moment you tell the second one “another AI says X”, you no longer have two opinions, you have one reacting to another.
A judge with a method. Someone reads every answer and compares them on substance, not on style: what’s verifiable, where they genuinely agree, where they materially differ. With explicit rules against the obvious biases — don’t reward the longest answer, nor the most self-assured, nor any particular provider.
The same standard every time. This is the hardest to see and the one that matters most. The jury rotates; the judge doesn’t. It’s a fixed agent with the same rubric on every question, so a confidence of 85 means the same thing today as next week.
What a jury tells you that one answer can’t
That they agree. When several AIs from different providers converge, that’s real information about how solid the answer is. None of them can tell you that about itself.
That they DON’T agree. This is the most valuable and least intuitive part. Disagreement isn’t a system failure: it’s the system warning you that your question depends on something you didn’t say, or that it has no automatic answer.
It happened to us in testing: we asked about a rent increase and some AIs said the landlord could, another said he couldn’t. None was lying — each quietly assumed a different scenario. With one, you’d have taken an assumption dressed as an answer.
How much to trust it. A number not produced by the same model that answered, but reflecting how much independent answers converged and how much of what they say is checkable.
What it doesn’t fix
Worth saying plainly: if all three AIs are wrong in the same way, the jury is wrong. Agreement reduces the risk of invention, it doesn’t remove it — some errors are shared because they come from the same place.
And for questions with a correct, checkable answer, a jury is overspending: if you can verify it at the source, verify it and move on.
Where it makes sense
On what can’t be checked and is expensive to get wrong: a decision with conditions, a document that needs interpreting, a choice that years depend on.
There, the difference between four answers on screen and a verdict with a confidence level isn’t convenience. It’s the difference between having information and having a decision made.
In one line
A comparator shows you the answers. A jury judges them by the same standard every time and tells you how much to trust the result.
The hard part was never asking several. It was comparing them well, every time.