The model was never the bottleneck
Answer
What stops a stalled pilot is a short list of missing mechanisms, and the model is not on it. Capability is close to universal: around 88% of firms use AI in at least one function, and around 60% report minimal or no value from it. What stops a workflow is the same short list every time: documents worse than the sample, exceptions with nowhere to go, nothing writing back to the system of record, and nobody having agreed what good enough means. None of the four is solved by a better model, which is why the choice of model is a procurement decision rather than the design.
TL;DR
- Trying a better model feels like progress, which is what makes it expensive. Six weeks later the new model is in place and the workflow is still not in production.
- Capability is not the shortage. Around 88% of firms use AI in at least one function and around 60% report minimal or no value from it, both from large global surveys.
- Four things stop a pilot, and none is the model: the documents are worse than the sample, the exceptions have nowhere to go, nothing writes back, and nobody agreed what good enough means.
- Model-neutral follows from that. If the model is not the constraint, which model you use is a procurement decision: price, latency, data terms, and what your provider agreement already allows.
- The remaining work is unglamorous and specific to you: agree the bar and the evaluation set first, work on your real documents, gate anything that writes, then keep measuring.
Every stalled pilot arrives at the same meeting
Every stalled AI pilot arrives at the same meeting. The demo worked. Someone senior saw it and liked it. Then months passed, and the thing is still not doing any real work. And almost always, the first idea in the room is to try a better model.
It is the wrong instinct, and it is an expensive one, because it feels like progress. Swapping models is quick, it is measurable in a benchmark, and it lets everyone avoid the harder conversation. Six weeks later the new model is in place and the workflow is still not in production, because the model was never what was stopping it.
The capability is real. It is also not the shortage
Start with what is not in dispute. The capability is real. Today's models read a messy invoice, follow a multi-step instruction, and write a defensible summary better than the tools most firms had two years ago, and they keep improving. If raw capability were the constraint, the last three years would have produced a wave of quietly automated back offices.
That is not what happened. Around 88% of firms now use AI in at least one function,01 and around 60% report minimal or no value from it.02 Both come from large global surveys, and read together they say something specific: the shortage is not access to capability. Nearly everyone has that. The shortage is the ability to get a working demo to survive contact with a real business.
So what actually stops it?
In every stalled pilot we have looked at, it is always the same short list, and none of it is about the model.
The documents are worse than the sample
The demo ran on ten clean PDFs. Production gets a scanned form someone photographed at an angle, a fifty-page report with the number buried in a table, and a supplier who changed their layout last month.
The exceptions have nowhere to go
A workflow that handles 90% of cases and drops the rest on the floor is not a working workflow. Someone has to see the other 10%, in a queue, with the context needed to decide.
Nothing writes back
Reading a document is the easy half. Getting the result into the system of record, with the right permissions, without creating a duplicate, without breaking when the API times out, is the half that takes the time.
Nobody agreed what "good enough" means
This is the quiet one, and it is the most common reason a pilot never ships. If nobody has said which metric matters, on which set of real cases, at what threshold, then no one can ever sign the release. So the project stays in evaluation forever, which looks like caution and is actually the absence of a decision.
Why we are model-neutral
Not one of the four is solved by a better model; all four are solved by engineering, and that is why we are model-neutral. It is worth being precise about what the word means, because it is not a hedge and not a refusal to have an opinion. It follows from the thesis: if the model is not the constraint, then which model you use is a procurement decision, and procurement runs on a checklist, not a leaderboard. The checklist is short: the price at your volume, the latency at your concurrency, the data and retention terms, the region your data may sit in, and what your existing provider agreement already covers. The finalists then face the only benchmark that matters, a held-out set of your own documents with a bar someone signed, and the cheapest model that clears the bar wins.
The position comes with its exceptions stated, because a neutrality with no counterexamples is a slogan. Very long documents, languages a given model reads poorly, and domains where one vendor's document reading measurably wins on your own held-out set are all real cases where the model choice does decide. The method does not change in those cases: measure the candidates on your documents and let the bar pick. And when the market moves, the swap is contained by design, because the part we build, the queue, the write-back, and the measurement, does not rest on the choice. Changing the model is a re-run of the bar, not a rebuild.
The last mile is the job
The uncomfortable part of this position is that the remaining work is unglamorous and specific to you. There is no product to buy that knows your exception rules. That is the whole argument: the last mile is not a small finishing task after the clever part. It is the job. The five mechanisms that make up that last mile, and the order to retrofit them into a pilot that already stalled, are the subject of why your AI pilot did not reach production.
So we start where the failure actually happens, in the order the Approach page sets out. Agree the bar and the evaluation set before any build. Work on your real documents, in your tenant. Put a human gate in front of anything that writes to a system of record. Then measure against the bar you agreed, and keep measuring, because production is not a launch date. It is a standing commitment.
Sources
Every load-bearing claim above, with its source and the date we checked it.
| Source | Reference | Accessed |
|---|---|---|
| 01 McKinsey, The state of AI in 2025 (5 November 2025), basis verified | https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai (opens in a new tab) | |
| 02 BCG, The Widening AI Value Gap, September 2025, basis verified | https://media-publications.bcg.com/The-Widening-AI-Value-Gap-Sept-2025.pdf (opens in a new tab) |
Frequently asked questions
- Will a better model get a stalled AI pilot into production?
- Usually not, because the model was rarely what stopped it. Swapping models is quick and measurable in a benchmark, so it feels like progress, but six weeks later the workflow is still not in production. The four things that actually stop it are the real documents, the exception path, the write-back, and the missing definition of good enough.
- How should we choose which model to use, then?
- As a procurement decision, with a checklist rather than a leaderboard: the price at your volume, the latency at your concurrency, the data and retention terms, the region your data may sit in, and what your existing provider agreement already covers. Then hold the finalists to the only benchmark that matters, your own held-out documents, and pick the cheapest one that clears your bar.
- What does model-neutral mean in practice?
- It is not a hedge and not a refusal to have an opinion; it follows from the argument. If the model is not the constraint, then which model you use is a procurement decision: price, latency, data terms, and what your provider agreement already allows. We pick one with you and change it when the market changes, because the part we build does not rest on that choice.
- Are there cases where the model choice does decide it?
- Yes, and naming them is what makes model-neutral a position rather than a hedge. Very long documents, languages a given model reads poorly, and domains where one vendor's document reading measurably wins on your own held-out set are all real cases. The way you find out is the same either way: measure the candidates on your documents, not on a leaderboard, and let the bar decide.