Ask most companies which AI model they use and you'll hear a trophy name, whatever topped the headlines that month. Ask why, and the answer is some version of "it's the smartest one." That answer costs real money, and it survives almost no scrutiny.
Because "smartest" answers a question nobody in your business is asking. A leaderboard measures a model's best day: the hardest problem, the cleanest prompt, everything going right. Your business doesn't run on anyone's best day. It runs on the two-hundredth similar email, the messy spreadsheet, the instruction you typed in a hurry at 4pm. What matters there is a different trait entirely: when the model has a bad moment, how often does that moment land on your work?
Hype rewards the peak. You pay for the average.
A model is a hire, so assess it like one
The useful mental shift is to stop shopping for software and start thinking about hiring. You would never hire someone because they gave one dazzling interview answer. You'd want to know how they behave on an ordinary Tuesday, how often they get things wrong, and what they cost.
The same questions apply to a model:
- How reliable is it at your actual work? Not on benchmark problems. On the tasks you'd hand it, with your level of messiness, over weeks of use. Reliability at your work is the only number that matters, and it can be very different from the leaderboard.
- What happens when it's wrong? A reviewed email draft that's off-tone costs you a minute. A number in a proposal that nobody catches can cost you a client. The same model can be a great hire for one task and a liability for another.
- What does it cost to do the same work? This is where honest math matters more than intuition, because the answer genuinely goes both ways. For high-volume, short tasks, a subscription model can cost a fraction of a salary. But a model working through complex work burns tokens constantly, and at serious volumes the cost to match what an experienced person does can rise past that person's salary. Sometimes the model is cheaper. Sometimes the human is. You only know by doing the math for your task, not by assuming.
That last point is worth dwelling on, because a lot of the current conversation skips it. It is tempting to declare that AI is obviously cheaper than an employee, or that it's obviously not as good. Both declarations are wrong as general claims. Whether a model outperforms a person at a given job depends on the work, on the person, and on how the work is set up. A careful senior employee with good judgment will beat a mediocre model at decisions with big consequences. A tireless model will beat an overloaded employee at a thousand repetitive drafts. The honest position is not "models are better" or "people are better." It's that each case needs its own assessment.
Running the assessment
When I help a company think through this, the sequence looks like the diligence you'd run before any important hire:
- Name the task and its blast radius. What actually happens when the output is wrong? Drafting and research usually have a cheap failure mode. Anything that reaches a customer, a regulator, or a contract does not.
- Decide the error rate you can live with. Every role in your company tolerates some mistakes; nobody demands flawless first drafts. Write down what "good enough" means before you look at a single model.
- Test candidates on your work, not theirs. Give each model a batch of your real tasks. Count errors. This takes an afternoon and it replaces weeks of guessing based on marketing.
- Do the cost comparison honestly. Put the model's monthly cost for the workload next to what you'd pay a person for the same output, including your time supervising the model. Accept that the answer may favor the human, especially for complex judgment work at moderate volume.
- Match the model to the consequence, not the trophy. For long stretches of work nobody watches, the odds of at least one failure climb fast, so reliability earns its premium. For supervised, reviewed tasks, a modest model is often all the hire you need.
- Re-run it every quarter. Today's mid-tier models were last year's frontier. A hire decision you made six months ago may be stale today, and the cheap option you rejected may now clear your bar.
None of this requires technical depth. It requires the same discipline you already apply to people decisions: define the job, test against the job, compare the full cost, and decide per case.
The trap is picking one answer for everything, whichever answer that is. Companies that put the trophy model on every task burn money. Companies that dismiss AI entirely because one test disappointed them leave real value on the table. The companies getting value are the ones assessing task by task, with numbers instead of headlines.
This is the same starting point I recommend everywhere: the workflow and its tolerance for error come first, the tool second. I cover the workflow side in the five questions to ask before automating anything.
If you're weighing which tasks deserve a premium model and which are burning money on one, that's exactly the kind of assessment I run with leadership teams: book a call and let's put your workflows through the same math.


