Can you trust AI to make decisions? When yes, and when not
“Should I trust what the AI just told me?” It’s a good question and it rarely gets an honest answer. Enthusiasts say yes to everything; sceptics say no to everything. Neither helps when you have a decision sitting in front of you this afternoon.
The useful answer is that it depends on the kind of decision, and the line can be drawn fairly clearly.
Three kinds of decision
Where AI is simply good. Decisions built on public information, where no private fact of yours is decisive and the cost of being wrong is low or reversible. What to pack, how to structure a difficult email, what to ask in an interview, how something you don’t understand actually works. No caveats needed here: use it.
Where it helps, but doesn’t decide. Important, nuanced decisions where the missing facts are yours alone: changing jobs, accepting an offer, a large purchase, a hard conversation. Here AI is genuinely useful for organising the problem — what’s in play, what you’re forgetting, what questions you should be asking — but the final call has a part you can’t delegate, because it rests on things you haven’t told it and probably couldn’t articulate.
Where you shouldn’t use it. Health. Personal crises. Anything with binding legal consequences or your own money at stake. Not because AI talks nonsense, but because the cost of an error isn’t recoverable, and because in those situations there are professionals whose job is exactly that — including being accountable for what they tell you.
That last part isn’t boilerplate to cover ourselves: it’s The Judge’s actual policy. On a health question we don’t arbitrate, we refer. We’d rather not answer than answer something that merely sounds right.
The three things that go wrong
If you keep three ideas about why to stay sceptical, make it these.
An AI never says it doesn’t know. It produces the most probable answer, not the truest one. Most of the time those coincide. When they don’t, nothing visible distinguishes them: the same confident tone covers what it knows cold and what it’s filling in.
The same question gives different answers. Reorder the sentences, ask tomorrow, or ask a different AI, and you can get different advice. If your decision depends on which one you happened to get, you weren’t deciding on information — you were tossing a coin.
It will agree with you. This is the hardest to catch because it feels good. If you phrase the question so your preference shows — and it almost always shows — the answer will lean that way. An AI has no incentive to push back.
How to use it without getting played
Four habits that change the outcome more than any phrasing trick:
Ask against yourself. Instead of “is this offer a good idea?”, try “what are the strongest reasons to turn this offer down?”. You’ll read things the first framing would never have surfaced.
Separate fact from judgement. Facts — figures, dates, regulations — get verified against a source, not argued with the AI. Reasoning you can evaluate yourself by reading it.
Cross-check when it matters. Ask two different AIs. If they agree, you have a real reason to trust it. If they disagree, you’ve just found the exact spot to think twice about.
Notice whether it hedges. A genuinely expert answer almost always contains a “that depends on…”. A flat, confident answer to a question that isn’t flat is a warning sign, not a sign of competence.
Trust isn’t a yes or a no
The word “trust” does damage here, because it sounds like a switch: either you trust it or you don’t. In practice trust has degrees, and the honest thing is to say which one you’re dealing with.
That’s why every verdict from The Judge carries an explicit confidence level. 91% and 68% read differently and should be acted on differently: at 68 the answer is still useful, but it’s telling you to look closely and that the last word is yours.
It’s also why we flag when the jury didn’t agree. That’s the information a system optimising to look competent would quietly hide.
What a jury actually fixes
No tool will remove uncertainty when a matter is genuinely uncertain. Promising that would be a lie.
What can be removed is the arbitrariness — your answer depending on which AI took the question, how you phrased the sentence, or what day it was. That’s where there’s a real difference between asking one and convening a jury with a judge who applies the same yardstick every time.
A loose opinion and an opinion backed by several aren’t worth the same, even when they say the same thing.
See how it works, or try it on whatever’s on your desk right now.