Skip to content
evals

Is 95% accuracy good enough, and who signs it off?

TheFrontierForgePublished Updated

Answer

No percentage is good enough on its own. An accuracy bar is a percentage plus a unit, a held-out set of the buyer's own cases, a stated cost for each of the two ways the system can be wrong, and the name of the person who signs it. Ninety-five percent per field on a twelve-field document and 95 percent per document are different systems, and only one of them has been priced.

TL;DR

  • No percentage is good enough on its own; an accuracy bar is a percentage attached to a unit, measured on your own cases, with a ceiling on the error class that reaches a customer, and a name against it.
  • A per-field bar and a per-document bar are different systems: at 95 percent per field, a document with twelve fields is fully correct only about 54 percent of the time.
  • The two error classes are not symmetric. Price a false accept and a false reject separately, and cap the one that reaches a customer or your system of record.
  • The answer key is written by the domain owner who can price the error, an underwriter or a CPA, not the engineer who builds the pipeline and not the vendor.
  • No external standard sets a testable bar before December 2027, so the contract is where the bar lives; payment is staged and the final tranche is tied to the agreed bar.

Is 95 percent accuracy good enough?

Ninety-five percent accuracy is not good enough on its own, and no other single percentage is either. A bare percentage in a contract is unenforceable, because both sides can satisfy it while measuring different things: the vendor counts correct fields, the buyer counts correct documents, and the same 95 describes two systems that behave nothing alike. An accuracy bar that means something is a percentage attached to a unit, measured on a held-out set of your own cases, with a separate ceiling on the one error class that reaches a customer or your system of record, and signed before the build by a person who can price the mistake. Most stalled pilots never had that bar, which is one reason pilots stall long after the model works.


Ninety-five percent of what, exactly?

"Ninety-five percent of what" is the question a threshold has to survive before it belongs in a contract. A number on its own has no unit, and per field, per document, or per decision are three different measurements that a single system can pass on one and fail on another. So the bar carries four things the percentage alone cannot: a unit that says what is being counted, a held-out set of the buyer's own cases that says where it was counted, a stated cost for each of the two ways the system can be wrong, and the name of the person who signs it. Drop any one and the remaining number can be met by a system nobody would ship.


Why do a per-field bar and a per-document bar produce different systems?

A per-field bar and a per-document bar produce different systems because field errors multiply across a document. Take a document with n fields and a per-field accuracy of p; the share of documents in which every field is correct is p raised to the n. At p of 0.95 and n of twelve, that is about 0.54, so roughly 46 percent of documents carry at least one wrong field. A per-document bar of 95 percent, by contrast, allows only 5 percent defective documents. The standard that sounds identical across a table is about nine times weaker in production, and the figure in the pilot deck was never the figure the room thought it had agreed.


What are the two ways the system can be wrong, and how do you price them?

The system can be wrong in two ways, and they do not cost the same. A false accept passes a wrong value through as correct, where it reaches a customer or writes into the system of record; a false reject flags a correct value for review, where it costs reviewer minutes and nothing else. On WorkBench, a sandbox benchmark revised on 1 July 2026, the best measured agent completes 97.7 percent of tasks and still takes an unintended harmful action, such as emailing the wrong person, on 1.9 percent of them.01 Completion and harm are two separate numbers, and a single headline percentage hides the one that decides everything. The authors add their own ceiling: any model trained after 2024 may have seen the benchmark, so the figure is an upper bound rather than a promise about your data.

Price the two classes separately before you set the bar, from counts on your own held-out set:

QuantitySymbolWhere the number comes from
False accepts per hundred itemsacounted on the held-out set
Cost of one false acceptC_adownstream correction or customer remediation
False rejects per hundred itemsrcounted on the held-out set
Cost of one false rejectC_rreviewer minutes to clear a good case

The class with the larger product of rate and cost is the one the ceiling has to cap. The cases that would breach it are the cases that route to a person.


What does the bar look like on a claims field and an invoice field?

The bar looks different on a claims field and an invoice field because the misses land in different places. On an auto claim, a license plate or policy number extracted from a police report can be checked in code against the vehicle record or the policy administration system, so a wrong value is caught before it writes; the bar sets a high per-field rate and a near-zero ceiling on those identity fields precisely because the cross-check makes the ceiling cheap to hold. On an accounts payable invoice, the total can be checked against the purchase order and the vendor bank detail against the supplier master, but a wrong bank detail that clears becomes an irreversible payment, so its ceiling is stricter than the total's. The same headline 95 percent is acceptable when the misses fall on a field a second source catches and unacceptable when they fall on a field that reaches money, which is why the unit and the ceiling, not the percentage, decide whether the system is safe.


Whose documents is the bar measured on, and who writes the answer key?

The bar is measured on the buyer's own cases, with the ugly ones left in, because a benchmark measures someone else's documents under someone else's definition of correct. That held-out set needs an answer key, and the answer key is written by the domain owner, not the engineer: an underwriter for policy exclusions, a CPA for a tax position. An engineer can build the pipeline and cannot write its answer key, because the key encodes judgments about coverage and materiality that live in a specialist's head rather than in the codebase. Let engineers invent the gold set for a domain they do not understand and the number that comes back is confident and meaningless.


Who signs the bar, and why is it not the engineer and not the vendor?

The person who signs the bar is the person who can price the error, which is neither the engineer who builds the pipeline nor the vendor who sells the model. The engineer optimizes toward whatever metric is written down and cannot say what a wrong coverage decision costs the business; the vendor's benchmark describes its model on its data and stops at your door. The signer is the owner who carries the consequence of a wrong value: the one who answers for a mispaid invoice or a wrongly declined claim. A bar signed by anyone else is a number without an owner, and a number without an owner is not defended when it slips.


What do you hand an auditor to show the bar was held?

You hand an auditor the record a bare percentage cannot produce: the signed bar with its unit, the held-out set and its size, the pass rate for each error class, the agreement between any model judge and the human labels, and the date and model version the measurement was taken against. A percentage on a slide is not auditable because it does not say what was measured, on whose cases, or when. The measurement envelope is the difference between showing an auditor the bar was held and asking one to take your word for it.


Is there a standard that sets the bar for you?

No external standard will set the accuracy bar for you before December 2027, so the contract is currently the only place a testable bar can live. Two enacted measures shape the record you will have to keep, and neither states a number.

The regulatory floor
MeasureDatesWhat it bears on
Colorado SB 26-189Signed 14 May 2026; applies to consequential decisions made on or after 1 January 2027A plain-language account of an adverse automated decision within thirty days, a process to request the system name and version, and an opportunity for meaningful human review and reconsideration, to the extent commercially reasonable
EU AI Act, digital omnibusParliament vote 16 June 2026; high-risk obligations apply from 2 December 2027 for stand-alone systems and from 2 August 2028 for embedded safety componentsA postponement to allow the necessary standards and support measures to be put in place; no numeric accuracy bar is set

Two enacted measures, dated. Both create obligations around automated decisions; neither defines a testable accuracy bar.

Colorado's SB 26-189 will require, for decisions made on or after 1 January 2027, a plain-language account of an adverse automated decision within thirty days and an opportunity for meaningful human review and reconsideration, to the extent commercially reasonable.02 The European Parliament voted on 16 June 2026 to postpone the AI Act's high-risk obligations to 2 December 2027 and 2 August 2028.03 Our reading of that record, stated as ours: no harmonised standard defines a testable bar before those dates, so a written bar you can show you held is what you can defend in the meantime.


What happens contractually if the finished system misses the bar?

What happens if the finished system misses the bar is written into the contract before the build, not discovered after it. The mechanism is staged payment: the work is priced as a fixed amount against the workflow, paid in stages, with the final tranche tied to the system clearing the agreed bar on the held-out set. The gate decides when the system deploys, not when the invoice falls due. Pricing the workflow rather than the model also puts the risk of a shifting token count or a swapped model on us instead of on your forecast. This is the shape of the Ship engagement: one workflow into production, on your stack, with the bar as the thing the final payment turns on.


When is a workflow not worth setting a bar for?

A workflow is not worth setting a bar for when the volume is low and a false accept costs almost nothing, or when no second source exists to check the consequential fields and every output would route to a person anyway. In the first case the bar costs more to define than the errors cost to absorb; in the second there is no automation to protect, only a queue with a model in front of it. Naming those cases is part of the work, because a bar defended on a workflow that should never have been automated is effort spent proving the wrong thing.

Two questions close the file, and both take about thirty seconds. Who signed your accuracy bar, and on whose documents was it measured? A reader who cannot answer either has just diagnosed the pilot. The readiness check at /assessment is where that signature gets obtained on a real document set, before a line of the build is written.

Sources

Every load-bearing claim above, with its source and the date we checked it.

SourceReferenceAccessed
01 WorkBench 2026 revisit (arXiv 2606.13715, v2), basis verifiedhttps://arxiv.org/abs/2606.13715 (opens in a new tab)
02 Colorado SB 26-189, basis verifiedhttps://leg.colorado.gov/bills/sb26-189 (opens in a new tab)
03 European Parliament, AI Act digital omnibus vote, basis verifiedhttps://www.europarl.europa.eu/news/en/press-room/20260611IPR45207/ai-act-ep-approves-simplification-measures-and-nudifier-app-ban (opens in a new tab)

Frequently asked questions

Is 95 percent accuracy good enough for an AI workflow?
No percentage is good enough on its own, because a bare number can be satisfied while each party measures something different. Good enough is a percentage attached to a unit, measured on a held-out set of your own cases, with a separate ceiling on the error class that reaches a customer or your system of record, signed before the build.
Who should sign off on the accuracy threshold?
The person who can price the error signs it, which is neither the engineer who builds the pipeline nor the vendor who sells the model. In a regulated domain that is the underwriter, the controller, or the CPA who carries the consequence of a wrong value and can adjudicate the answer key.
What is the difference between a per-field and a per-document accuracy bar?
A per-field bar sets the rate for each field; a per-document bar sets it for the whole document, and field errors multiply. At 95 percent per field across twelve fields, only about 54 percent of documents are fully correct, against 95 percent under a per-document bar.
Is there a legal standard that sets the accuracy bar?
No external standard defines a testable accuracy bar before December 2027. Colorado's SB 26-189 requires an explanation of an adverse automated decision within thirty days from 1 January 2027, and the EU has postponed its high-risk obligations to 2 December 2027, but neither sets a number, so the contract is where the bar lives.

If this maps to a workflow you want in production, the workflow review is the place to start.

  1. Book a workflow review
  2. Commission a Production Readiness Assessment
  3. Commission the build