Skip to content
AI Driven Solutions
All insights

AI7 min read

“95% accurate” is not a number you can run a business on

What accuracy claims in document AI actually measure, and the questions to ask before you trust one.

Every document extraction tool quotes an accuracy figure. Almost none of them define it the same way, and the definition matters far more than the number.

Accuracy of what, exactly

A vendor quoting 95% might mean 95% of individual fields extracted correctly, or 95% of documents fully correct, or 95% of characters recognised. On a twelve-field invoice, 95% per field means roughly half of your documents contain at least one error. Those are wildly different businesses to run.

Ask which fields, on whose documents

Extraction accuracy is not uniform. Invoice totals are easy; line items on a supplier who prints tables without borders are hard. And a figure measured on a public benchmark set says almost nothing about your post from your two hundred suppliers. We would agree a representative sample of your documents for training and evaluation, including difficult layouts and seasonal variation. The evaluation set must be kept separate from the training data.

Confidence is the useful number

The design decision that actually protects you isn't accuracy, it's what happens below a confidence threshold. Every extraction returns a confidence score per field. Set the threshold with the finance team, route anything under it to a human queue, and the failure mode becomes 'a person checks it' instead of 'a wrong number reaches the ledger'.

You're not buying accuracy. You're buying a system that knows when it isn't sure.

Measure drift, not just launch

A supplier redesigns their invoice template and your extraction quietly degrades. Report accuracy monthly against the exception queue's resolution outcomes, and you'll see it. Measure once at go-live and you'll find out from your auditor.

We'd rather show you than write about it.

Book a discovery call