Skip to content
document automation

Why your AI pilot did not reach production

TheFrontierForgePublished Updated

Answer

Your AI pilot did not reach production because it was built to prove a model works, not to survive an operation. What is missing is rarely accuracy; it is the five mechanisms production runs on: an accuracy bar agreed before the build, an exception queue, a human review gate, write-back into the system of record, and a named owner with an audit trail. Teams that treat these as scope from day one ship; teams that treat them as cleanup stall. Most stalled pilots can be finished by installing what was skipped, not by rebuilding the model.

TL;DR

  • Nearly every organization now uses AI; few get value from it. The constraint is implementation, not model access.
  • A pilot proves the model can do the task. Production asks what happens when it cannot: who catches it, who approves it, where the result lands, who answers for it.
  • The five mechanisms: an accuracy bar agreed before the build, an exception queue, a human review gate, write-back into the system of record, and a named owner with an audit trail.
  • None of the five requires a better model, and none appears on its own after go-live. They are scope, not cleanup.
  • A stalled pilot is usually finishable: keep the model, install the missing mechanisms, measure against the bar before go-live.

The numbers say implementation, not models

AI use is close to universal: 88% of organizations report regular use of AI in at least one function.01 Value is not: 60% of companies report minimal or no value from AI investment,02 and about 6% qualify as high performers on McKinsey's definition.01 Returns tell the same story. In Bain's Automation and AI Pathfinder Survey 2026 of 951 companies, 37% targeted cost reductions of 11 to 20%, and nearly 40% of those that measured outcomes landed at 0 to 10% instead.03

These are not model-quality numbers. The models did the demo fine. The gap opens after the demo, at the questions a pilot never has to answer: what happens to the case the system is unsure about, who approves an action before it becomes a customer's problem, where the result lands, and who answers for the system in an audit. A pilot with no answers to those questions was not almost done. The distance from demo to production is a set of mechanisms, and none of them is the model.


The five mechanisms

They run in the order a workflow meets them: the bar before the build, the queue and the gate during the work, the write-back where the result lands, the owner and the trail for every day after go-live.

1. An accuracy bar, agreed before the build

Before anything is built, the owner of the workflow signs a definition: what correct means for this work, field by field or decision by decision, and how it will be measured, on a held-out set of your own cases with the ugly ones left in. Not a benchmark score; a benchmark measures someone else's documents.

The bar changes behavior on both sides. The build team designs to clear it instead of to impress, and go-live becomes a milestone the system passes instead of a mood the room is in. Without it, the pilot is judged by anecdote, and anecdote always loses: one bad output in a demo outweighs everything the system got right, because nobody agreed in advance how much wrong was acceptable.

2. An exception queue

No extraction pipeline clears every case, and a system that pretends otherwise fails silently, which is the expensive way. The mechanism is a queue: the system grades its own confidence, and anything under the threshold routes to a person, with the document, the extracted fields, and the reason for doubt on one screen. The exception rate becomes a number someone watches: measured, not assumed.

Exceptions are expected, not exceptional. A pilot without a queue treats every edge case as a surprise, and each surprise spends trust the demo earned. The same edge case, arriving in a queue, is routine work.

3. A human review gate

The queue catches what the system doubts; the gate catches what the business cannot afford to get wrong even when the system is confident. A payment, a regulatory filing, anything a customer sees: those pass a person before they happen, not in a report afterward, until the measured record justifies widening the gate.

A gate only holds if reviewing is faster than doing the task by hand. That means a real work surface: approve in one action, correct in place, and every correction feeds the eval set so the system improves exactly where it was wrong. A slow gate gets bypassed, and a bypassed gate is worse than none, because everyone upstream still believes it is there.

4. Write-back into the system of record

A pilot that ends in a spreadsheet has not automated the workflow; it has added a step, because a person still re-keys the result into the system the business runs on. The mechanism is integration: clean records written into your ERP, claims platform, or CRM, under your existing permissions, with validation at the boundary and the timeouts, retries, and fallbacks that survive a bad day.

This is usually the largest piece of the build and the first one a pilot defers. The model is the easy part; the write-back is the work, and until it exists the workflow has not changed for anyone who runs it.

5. A named owner and an audit trail

Production systems answer questions long after the fact: what the system did with this document, who approved that write, what the exception rate was in March. The trail answers them: every action logged with what the system saw, what it did, and who signed off, kept where an auditor can reach it without an engineer in the room.

The owner is the other half. One named person owns the bar, the exception rate, and the cost per document the way a manager owns a team's output. A pilot without an owner is orphaned at the first reorg; a system without a trail is stopped by the first audit. The two keep each other alive: the owner needs the trail to manage, and the trail needs an owner to matter.


Finishing a stalled pilot

If your pilot is stalled, the model you already have is probably not the problem, and the way to find out is on this list. Agree the bar and measure against it; stand up the queue and the gate; wire the write-back; name the owner and turn on the trail. That is finishing, not rebuilding, and it is the shape of the Ship engagement: one workflow, into production, on your stack. How we take a workflow to production, stage by stage, with the gate each stage must pass, is on the Approach page. Once it is live, keeping it right is a discipline of its own; that is Operate.

Sources

Every load-bearing claim above, with its source and the date we checked it.

SourceReferenceAccessed
01 McKinsey State of AI 2025, basis verifiedhttps://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai (opens in a new tab)
02 BCG, The Widening AI Value Gap, September 2025, basis verifiedhttps://media-publications.bcg.com/The-Widening-AI-Value-Gap-Sept-2025.pdf (opens in a new tab)
03 Bain and Company, 2026, basis verifiedhttps://www.bain.com/insights/your-ai-budget-is-growing-your-returns-arent-heres-why/ (opens in a new tab)

Frequently asked questions

Was the model the problem?
Usually not. If the model misses the agreed bar on your own documents, that surfaces as soon as measurement starts, and swapping models is a contained decision. What stalls pilots is the absence of the bar, the queue, the gate, the write-back, and the owner, and no model upgrade supplies those.
Can a stalled pilot be finished, or does it need a rebuild?
Most stalled pilots are finishable. We instrument what exists, measure it against a bar the owner signs, and install the mechanisms it skipped; a rebuild is a last resort you would see coming in the measurements.
What does an accuracy bar look like in practice?
A short written statement per field or per decision: what correct means, how it is measured, and on which held-out set of your real cases. The number is yours to set with us; what we commit to is that it is measured before go-live, not asserted after it.
What does it cost to find out?
The workflow review is free: thirty minutes, and you leave with a one-page scope and a fixed price for your case, whether the starting point is a stalled pilot or a workflow that was never automated.

If this maps to a workflow you want in production, the workflow review is the place to start.

  1. Book a workflow review
  2. Commission a Production Readiness Assessment
  3. Commission the build