AI Data Extraction Accuracy SLA: What to Demand From a Vendor
You have three proposals on your desk and all three say 99 percent accuracy. One is an IDP platform, one is an offshore keying shop, one is an AI-first startup. None of them tells you 99 percent of what, measured on which documents, checked by whom, or what happens when the number comes in lower. That gap is where the money leaks.
Short Answer
An accuracy SLA is meaningless unless it names the denominator. Demand field-level accuracy on a named field list, measured against a double-checked ground truth on a holdout sample of your own documents, reported per field rather than as one headline number. Then attach a remedy: free rework on any field below its threshold, with a defined turnaround, and a termination right if the threshold is missed twice in a row.
The Five Denominators, and Why Only Two Matter
The same run can be described as 99.4 percent accurate or 55 percent accurate depending on what you divide by. Vendors are not always lying. They are choosing the flattering fraction.
| Metric | What it divides by | Typical vendor use | What it actually tells you |
|---|---|---|---|
| Character accuracy (CER) | Individual characters recognised | The source of most 99 percent headline claims | Almost nothing. One wrong digit in an invoice number is a total field failure at 99.9 percent character accuracy |
| Field accuracy | Named fields extracted correctly | Rarely volunteered, usually available on request | The number you should contract on. Report it per field, not as an average |
| Document accuracy | Documents with zero errors in any field | Almost never quoted | Your true touchless rate. At 97 percent field accuracy across 20 fields, roughly 54 percent of documents are clean end to end (0.97^20) |
| Line-item accuracy | Rows in tables, not header fields | Quietly excluded from headline numbers | Where nearly all the failures live. Contract on this separately |
| Straight-through processing rate | Documents needing no human touch | Used as a proxy for accuracy | An operations metric, not an accuracy metric. It rises when you loosen review rules |
Two of these belong in a contract: field accuracy per named field, and line-item accuracy where tables exist. Character accuracy belongs nowhere near a commercial agreement. Document accuracy is useful as a sanity check, because the compounding arithmetic exposes any headline claim that cannot survive twenty fields. Ask which one the 99 percent refers to, in writing.
Have a messy PDF? Upload 1-3 sample pages and we will tell you if it is clean, OCR-heavy, or needs human QA.
Step One: Write the Field List Before You Ask for a Price
No vendor can commit to accuracy on an undefined scope. Before the RFP goes out, produce a field register with four columns: field name, source location, format rule, and criticality tier. Three tiers are enough.
- Tier 1, financial and identifying. Invoice total, tax amount, account number, IBAN, policy number, date of service. An error here causes a payment or compliance problem. Threshold at or above 99.5 percent.
- Tier 2, operational. Vendor name, PO reference, currency, due date. An error causes downstream rework. Threshold around 98 to 99 percent.
- Tier 3, descriptive. Line descriptions, notes, free text. Annoying, not expensive. Threshold around 95 percent, or excluded from the SLA.
A single blended threshold lets a vendor pass by nailing easy descriptive fields while missing totals. Tiering removes that escape route.
Step Two: Define What Counts as an Error
This is the clause most buyers skip and most disputes turn on. Agree in writing on each failure mode.
- Wrong value. An error. Uncontroversial.
- Missing value where one exists on the page. An error. Some vendors argue nulls are abstentions rather than mistakes. Reject that.
- Value returned where the field is genuinely absent. A hallucinated value is the worst failure mode because it passes review. Count it, and consider double-weighting it on Tier 1 fields.
- Format deviation. 03/04/2026 when the spec says ISO dates. Decide up front whether this is an error or a normalisation task the vendor owes you for free.
- Partial match. "Acme Corp" versus "Acme Corporation Ltd". Fix the matching rule, exact or normalised, before measurement rather than after.
Step Three: Specify the Measurement Method
An accuracy number without a method is a marketing asset. The method needs five things.
A holdout sample of your documents. Give the vendor a tuning set to configure against and keep a separate holdout set they never see. Vendor benchmark decks run on curated corpora, and production performance on real mail-room input routinely lands well below them.
A representative mix. Include the long tail deliberately: photocopies, phone photos, rotated pages, faxes, handwriting, tables that split across a page break, and any legacy template still in circulation. If 8 percent of your monthly volume is bad scans, the sample should be about 8 percent bad scans.
A double-checked answer key. Two independent humans key the holdout set and disagreements get adjudicated. A single-pass answer key contains errors, and those errors show up as vendor failures or, worse, as vendor passes.
A defensible sample size. For a 99 percent threshold, a 50-document sample tells you almost nothing, because one error moves the measured rate by two points. Size the sample so a single error cannot swing the result past the threshold, and state that size in the contract.
A reporting cadence. Monthly per-field reporting on a rolling sample of live output, not a one-time acceptance test. Acceptance-test-only SLAs decay quietly as your document mix drifts.
Step Four: Attach a Remedy With Teeth
Service credits sized at a few percent of monthly fees are theatre. If 4,000 invoices come back with wrong totals, a 5 percent credit does not pay for the correction work. Ask for a remedy ladder instead.
| Miss | Remedy to require |
|---|---|
| Any Tier 1 field below threshold | Free re-extraction of the affected batch, corrected file delivered within a stated turnaround |
| Second consecutive monthly miss on the same field | Free rework plus a written root-cause note and a remediation plan with a date |
| Third consecutive miss, or any Tier 1 field below a hard floor | Termination for cause with no exit fee, and delivery of your data in an open format |
| Systemic miss traced to vendor process change | Rework plus credit against the affected month, not the next one |
Two more clauses are worth the negotiation. First, no silent model swaps: require notice before a change to the underlying model or pipeline, because an upgrade can raise average accuracy while regressing one field type. Second, your data is yours, it is not used for training, and it is deleted on a stated schedule.
Confidence Scores Are a Routing Tool, Not a Guarantee
Every vendor will show you confidence scores and a review queue. Treat those scores as an unvalidated input. Recent benchmarking of vision-language models on document extraction found calibration varies widely across models, from close to well-calibrated down to severely overconfident, and that the best way to produce a confidence signal differs by model. A model can be confidently wrong, and those errors sail past a threshold-based review gate untouched.
So test the gate. On your holdout set, measure what share of genuine errors the review trigger catches and what share of flagged items turn out to be fine. A gate that flags 40 percent of documents to hit its accuracy number is not automation, it is a manual process with an AI-shaped invoice.
Price on the same basis: cost per successfully processed document and cost per accepted field, with retries, failed jobs, premium modes, and your own review hours rolled in. A lower page rate that pushes a third of documents into your queue is usually the more expensive option.
When a Software Tool Beats a Managed Service
If your documents come from a small set of stable templates, volume is high and steady, and you have engineers who can own a pipeline, buy an IDP platform or build on a document AI API. Lower unit cost, full control, and an SLA you write yourself against your own eval set.
A managed service makes sense when the mix is messy and varied, volume is lumpy, the field list changes, or you do not want to run an extraction pipeline at all. DataConvertPro sits on that side: humans plus AI, delivering checked Excel and CSV, with the commitment made on the delivered file rather than on a model's benchmark score. Either way the checklist is the same. Name the fields, name the denominator, name the measurement, name the remedy.
Frequently Asked Questions
Is 99 percent accuracy good for AI data extraction?
It depends on the denominator and the field count. At 99 percent field accuracy across 20 fields, only about 82 percent of documents come out fully clean, so one in five still needs a human. On Tier 1 financial fields, 99 percent is often not good enough, because a wrong total on 1 invoice in 100 is real cash exposure. Set thresholds per tier rather than accepting one blended figure.
How large should the accuracy test sample be?
Large enough that one error cannot move the measured rate past your threshold. Testing a 99.5 percent commitment on a 100-document sample puts a single error at 99 percent, and the test cannot separate an unlucky batch from a real miss. Size the sample against the threshold you are enforcing, write that number into the contract, and re-measure monthly on live output rather than once at onboarding.
What if the vendor refuses to sign a field-level accuracy SLA?
That is informative rather than automatically disqualifying. Many good tool vendors will not sign one, because they cannot control your document quality. That is defensible for software you configure yourself. A managed service that will not commit at the field level is a different matter, since scope, method, and remedy are exactly what you are buying. Ask what they will commit to instead: turnaround, rework terms, or a paid pilot on your holdout set.
Should the pilot be paid?
Yes, if you want a serious answer. A free pilot gets a curated sample and the vendor's best engineer. A paid pilot on a representative holdout set, scored against a double-checked answer key, gives you a number you can put in the contract and gives the vendor a reason to tell you honestly which document types will be difficult.
Ready to Convert Your Documents?
Stop wasting time on manual PDF to Excel conversions. Get a free quote and learn how DataConvertPro can handle your document processing needs with AI-assisted extraction and human verification.