How to choose which AI to use for what
“Which is the best AI?” has no useful answer, and we covered that elsewhere. But the question that does matter — which one do I use for the thing in front of me — can be answered, and it doesn’t require following each month’s leaderboard.
The trick is to stop thinking in brands and start thinking in kinds of task.
Four kinds of task, and what each one needs
Tasks with a correct answer. A fact, a calculation, a translation, a spelling. What matters here is that it doesn’t invent. Any of the big ones will do, and what decides the outcome isn’t which you pick but whether you check the verifiable part.
Creative tasks. A text, a script, a name, a structure. There’s no correct answer, there are options. Here the useful move isn’t choosing well: it’s seeing two or three different versions and taking pieces from each. Picking a single tool is precisely what removes the value.
Reasoning tasks. A decision with conditions, an analysis, a problem with several variables. Here there are real differences between models — and here is also where relying on one is most dangerous: a long, wrong chain of reasoning reads exactly as well as a correct one.
Document tasks. A contract, a report, a tender. What decides this isn’t the model’s intelligence but whether it actually read the file — and that varies enormously, even among those that claim to accept them. We measured it in what happens when you give an AI a PDF.
The rule that saves you keeping up
Instead of maintaining a table of which model is best this month, use this:
If the answer can be checked, it doesn’t matter who gave it. If it can’t be checked, don’t trust just one.
A legal fact you verify at the source, and it doesn’t matter which AI told you. Advice on whether to accept a job offer can’t be verified anywhere — which is why one opinion, however good the brand, leaves you exactly as alone as before.
Mistakes when choosing
Choosing by the last comparison you read. Rankings shift every few months and usually measure things that aren’t your task.
Always using the same one out of habit. It’s comfortable, and it’s how you end up not knowing another did your thing better.
Choosing by price on expensive decisions. Saving pennies on the question that settles a three-year contract is a strange economy.
Asking an AI which AI is best. It will answer enthusiastically and has no way of knowing.
The alternative: don’t choose
There’s one case where the right answer is to choose none: when the decision matters and can’t be verified.
There, value doesn’t come from picking the right tool but from seeing several independent answers and looking at where they agree and where they don’t. Agreement between sources that never spoke says something none of them can say about itself; disagreement warns you your case has a wrinkle.
That’s exactly what The Judge is for: you don’t pick a model, several are convened and a fixed judge compares them on substance and gives you a confidence level. For anything checkable, keep using whatever you have to hand — nothing else is needed there.
In one line
Don’t look for the best AI. Look at what kind of task you have, and ask whether the answer can be checked.
That decides it well 90% of the time, and you never have to read another comparison.