Skip to content
llm integration

The model was never the bottleneck

TheFrontierForgePublished Updated

Answer

The model was never what stopped the pilot. Capability is close to universal: around 88% of firms use AI in at least one function, and around 60% report minimal or no value from it. What stops a workflow is the same short list every time: documents worse than the sample, exceptions with nowhere to go, nothing writing back to the system of record, and nobody having agreed what good enough means. None of the four is solved by a better model, which is why the choice of model is a procurement decision rather than the design.

TL;DR

  • Trying a better model feels like progress, which is what makes it expensive. Six weeks later the new model is in place and the workflow is still not in production.
  • Capability is not the shortage. Around 88% of firms use AI in at least one function and around 60% report minimal or no value from it, both from large global surveys.
  • Four things stop a pilot, and none is the model: the documents are worse than the sample, the exceptions have nowhere to go, nothing writes back, and nobody agreed what good enough means.
  • Model-neutral follows from that. If the model is not the constraint, which model you use is a procurement decision: price, latency, data terms, and what your provider agreement already allows.
  • The remaining work is unglamorous and specific to you: agree the bar and the evaluation set first, work on your real documents, gate anything that writes, then keep measuring.

Every stalled pilot arrives at the same meeting

Every stalled AI pilot arrives at the same meeting. The demo worked. Someone senior saw it and liked it. Then months passed, and the thing is still not doing any real work. And almost always, the first idea in the room is to try a better model.

It is the wrong instinct, and it is an expensive one, because it feels like progress. Swapping models is quick, it is measurable in a benchmark, and it lets everyone avoid the harder conversation. Six weeks later the new model is in place and the workflow is still not in production, because the model was never what was stopping it.


The capability is real. It is also not the shortage

Start with what is not in dispute. The capability is real. Today's models read a messy invoice, follow a multi-step instruction, and write a defensible summary better than the tools most firms had two years ago, and they keep improving. If raw capability were the constraint, the last three years would have produced a wave of quietly automated back offices.

That is not what happened. Around 88% of firms now use AI in at least one function,01 and around 60% report minimal or no value from it.02 Both come from large global surveys, and read together they say something specific: the shortage is not access to capability. Nearly everyone has that. The shortage is the ability to get a working demo to survive contact with a real business.


So what actually stops it?

In every stalled pilot we have looked at, it is always the same short list, and none of it is about the model.

The documents are worse than the sample

The demo ran on ten clean PDFs. Production gets a scanned form someone photographed at an angle, a fifty-page report with the number buried in a table, and a supplier who changed their layout last month.

The exceptions have nowhere to go

A workflow that handles 90% of cases and drops the rest on the floor is not a working workflow. Someone has to see the other 10%, in a queue, with the context needed to decide.

Nothing writes back

Reading a document is the easy half. Getting the result into the system of record, with the right permissions, without creating a duplicate, without breaking when the API times out, is the half that takes the time.

Nobody agreed what "good enough" means

This is the quiet one, and it is the most common reason a pilot never ships. If nobody has said which metric matters, on which set of real cases, at what threshold, then no one can ever sign the release. So the project stays in evaluation forever, which looks like caution and is actually the absence of a decision.


What those four have in common

Notice what those four have in common. Not one of them is solved by a better model. They are solved by engineering: by handling your real documents, building the review path, doing the integration, and agreeing a bar in writing before the build starts.


Why we are model-neutral

This is why we are model-neutral, and it is worth being precise about what that means. It is not a hedge and it is not a refusal to have an opinion. It follows from the thesis. If the model is not the constraint, then which model you use is a procurement decision: price, latency, data terms, and what your provider agreement already allows. We will pick one with you and we will change it when the market changes, because the part we build does not rest on that choice.


The last mile is the job

The uncomfortable part of this position is that the remaining work is unglamorous and specific to you. There is no product to buy that knows your exception rules. That is the whole argument: the last mile is not a small finishing task after the clever part. It is the job.

So we start where the failure actually happens, in the order the Approach page sets out. Agree the bar and the evaluation set before any build. Work on your real documents, in your tenant. Put a human gate in front of anything that writes to a system of record. Then measure against the bar you agreed, and keep measuring, because production is not a launch date. It is a standing commitment.

Sources

Every load-bearing claim above, with its source and the date we checked it.

SourceReferenceAccessed
01 McKinsey State of AI 2025, basis verifiedhttps://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai (opens in a new tab)
02 BCG, The Widening AI Value Gap, September 2025, basis verifiedhttps://media-publications.bcg.com/The-Widening-AI-Value-Gap-Sept-2025.pdf (opens in a new tab)

Frequently asked questions

Will a better model get a stalled AI pilot into production?
Usually not, because the model was rarely what stopped it. Swapping models is quick and measurable in a benchmark, so it feels like progress, but six weeks later the workflow is still not in production. The four things that actually stop it are the real documents, the exception path, the write-back, and the missing definition of good enough.
What actually stops an AI pilot from reaching production?
The same short list every time. Production documents are worse than the demo sample. The cases the system cannot handle have nowhere to go. Nothing writes the result back into the system of record. And nobody has said which metric matters, on which set of real cases, at what threshold, so no one can sign the release.
What does model-neutral mean in practice?
It is not a hedge and not a refusal to have an opinion; it follows from the argument. If the model is not the constraint, then which model you use is a procurement decision: price, latency, data terms, and what your provider agreement already allows. We pick one with you and change it when the market changes, because the part we build does not rest on that choice.
If the model is not the constraint, what is the work?
Engineering, and it is specific to you. Handling your real documents, building the review path for the cases that stop, doing the integration into the system of record, and agreeing a bar in writing before the build starts. There is no product to buy that knows your exception rules.

If this maps to a workflow you want in production, the workflow review is the place to start.

  1. Book a workflow review
  2. Commission a Production Readiness Assessment
  3. Commission the build