The Cheaper Model Was Right 97% of the Time. We Didn't Ship It.

Rahul Bhattacharya
Co-Founder and CTO, Adopt AI29 September 2026

Last week we ran a head-to-head between two models on one of the dullest jobs in accounting: reading a bank line like SQ *BLUE BOTTLE COFFEE 4417 OAKLAND CA and deciding which of 32 categories it belongs to.
One was a small, purpose-built decision model: it doesn't generate text at all; it picks an answer from a menu and returns a probability. The other was a small general-purpose LLM, the kind most AI bookkeeping tools run on without saying so.
The decision model beat the LLM on accuracy, ran five times faster, and cost a fraction per line.
We're not shipping it, and the reason matters more than the result if you're evaluating any AI for a set of books.
What the test was
Categorisation is a closed-menu decision. Given five fields (description, merchant, amount, currency, date), pick one of 32 groups (software_it, groceries, professional_services, …). One of those groups is an escape hatch: other_unclassified, which means I don't know, a person should look. Every proposal stays pending until a human confirms it. Nothing posts on its own.
We wrote 157 bank lines by hand and labelled them. 107 are clear from the descriptor alone. 50 are deliberately ambiguous, in five families a bookkeeper will recognise:
- bare bank memos:
CHECK 1061,WIRE OUT 20260903 TRN 0007731,ONLINE TRANSFER TO XXXXXX4471 - pass-through rails that hide the payee: PayPal, Venmo, Zelle, Wise, a Stripe transfer
- generic card-processor merchants:
SQ *M CHEN,SQ *THE STUDIO - marketplaces and warehouse clubs: Amazon Business, Costco, Sam's Club, Walgreens, eBay
- one brand, several businesses: Apple,
GOOGLE *SVCS, Microsoft, LinkedIn Premium, ADP fees
On those, the right answer is usually not to guess.
Then each model ran the full set five times, through the same gateway, so every number below is a mean with its range, not something we saw once.
The numbers
| Decision model | General LLM | |
|---|---|---|
| Agreement, clear lines | 97.1% | 94.7% (94.2–95.1) |
| Declined to answer on ambiguous lines (higher is better) | 22.0% | 38.8% (38.0–40.0) |
| Right when it did commit on an ambiguous line | 56.4% | 73.2% (71.0–74.2) |
| Ambiguous lines it got wrong and confident, out of 50 | 17.0 | 8.2 (8–9) |
| Same answer on every run | 99.4% of lines | 97.5% of lines |
| Wall clock, 157 lines | ~10 s | 51 s |
| Model calls | 157 | 7 |
Accuracy: the cheaper model wins. Ambiguous lines: it guesses more, and is wrong more when it does.
Read the first row and you'd ship the cheap one tomorrow. Read the next three and you'd stop.
On the 50 lines that were unclear, the cheaper model guessed on 39 of them and got 17 wrong: a third of the ambiguous set, confidently miscategorised, every single run. The LLM guessed on 31 and got 8 wrong, twice as many confident mistakes from the model with the higher accuracy score.
That gap is exactly what "accuracy" hides.
Why abstention is the metric
There are four things a categorisation model can do with a line:
| Commits | Abstains | |
|---|---|---|
| Right | Fine. Saves ~8 seconds. | Fine. Costs ~8 seconds. |
| Wrong | The expensive box. | Fine. Costs ~8 seconds. |
Three of the four boxes cost a bookkeeper a few seconds. The fourth, wrong but confident, costs something worse.
A confident wrong entry doesn't look wrong. It sits in the ledger with a confidence of 91. It rolls into the wrong P&L line, gets the wrong tax treatment, and is either caught in review (by someone who now trusts the tool less), caught by the auditor (worse), or never caught. A blank costs eight seconds. A confident wrong answer costs an hour if you're lucky and a restated month if you're not.
The accuracy number, the one every vendor leads with, is measuring the three cheap boxes together. The number that decides whether the tool belongs in a CPA's workflow is how often it lands in the fourth. On this test the more accurate model landed there twice as often.
Accuracy: the cheaper model wins. Abstention: it loses by two to one.
"Just raise the confidence threshold"
That's the first objection, and it's a fair one, so we checked. Both models return a confidence, and both are gated at 60 today. Raise the bar to 85, and a bookkeeper would see this:
| Ambiguous lines, threshold 60 → 85 | Decision model | General LLM |
|---|---|---|
| Declines to answer | 22% → 32% | 39% → 50% |
| Wrong and confident, out of 50 | 17 → 13.4 | 8.2 → 4.2 |
Confidence carries information on both models, but 90% of answers sit above 85, so a threshold barely reaches the mistakes.
Raising the threshold helps both, and the gap widens: three to one instead of two to one. The reason is in the calibration. Both models' confidence does carry information (the 85+ band is more often right than the 60–85 band on each), but more than 90% of all answers sit above 85, including the wrong ones. The decision model gave 15 wrong answers per run at 85 or higher. No threshold you'd set reaches those.
Take two things from that. A confidence score is only useful if you've looked at accuracy per band: ask any vendor for that table. And a threshold trims the edge of the problem; it doesn't change which model knows when to stop.
What "cheaper and faster" buys
The decision model is one call per line, ~375 ms each, run in parallel: the whole set in about ten seconds against fifty-one for the LLM. Per line it costs a fraction of a cent on either model. Both are rounding errors against the bookkeeper's hour.
One wrong entry that reaches the financials costs more than the entire year's model bill, on either model. Cost is a tiebreaker at best. It should never be the axis you choose on, and in most vendor comparisons, it's the only axis shown.
What it would take to ship
We're not writing the decision model off. It's fast enough to run inside the typing loop, cheap enough to run on every line instead of a sample, and it gave the same answer on 99.4% of lines across five runs; the LLM flipped on 2.5%. The bar to ship it is written down, and it's four numbers on at least 200 lines that a human decided, from at least two clients:
- Confident agreement at or above the LLM's. (Passes here.)
- Abstention on ambiguous lines no worse than the LLM's. (Fails here: 22% vs 39%.)
- The 85+ confidence band more often right than the 60–85 band. (Passes here: 88% vs 57%.)
- p95 latency under one second. (Passes: 584 ms.)
It passes three of four. The one it fails is the one that matters.
One caveat on everything above, because a test that isn't honest about its limits is worse than no test: these are 157 lines written by someone who knew the answer. That's a shape check: it catches a model that misreads descriptors or ignores the escape hatch. The real evaluation runs on lines a bookkeeper decided, where the disagreements are real judgement calls rather than ours. The numbers will move. The pattern held across five runs.
If you're evaluating AI for your books
Whatever the tool, ask for these five things. If a vendor can't produce them, they haven't measured them.
- Abstention rate on ambiguous lines, not accuracy alone. A model that never says "I don't know" is guessing somewhere.
- Wrong-but-confident, as its own number. It's the only error that costs real money.
- A calibration table. Accuracy per confidence band. If 90% of answers sit above the threshold, the threshold isn't protecting you.
- Ground truth from real human decisions, held out per client, not a fixture the vendor wrote.
- Cost last. It's a tiebreaker. The first four are the decision.
The cheapest categorisation model we've tested is right 97% of the time. The one we run is right 95%. We run the 95% one, because of what each does with the lines it shouldn't answer.
Numbers from an internal evaluation, 29 September 2026: 157 synthetic lines (50 ambiguous), five runs per model, both via the same gateway, confidence threshold 60. Scripts and fixture are committed in our repo; no client data was used.
Related Articles

The AI Controller: Job Description for a Role That Didn't Exist Two Years Ago


Will AI Replace Accountants? What the Data Actually Says

See it running on your workflows
Automate your accounting workflows

