Getting it live was chapter one. Keeping it right is the work.
One engagement, two modes: the build puts a workflow live; the retainer keeps it reliable, integrated, and governed, in your tenant.
What changes after go-live
Models are retired, document formats drift, volumes grow, and costs creep token by token. Each new integration widens what the system touches. None of this is failure; it is what production means. Operate is the standing form of the same engagement: the team that puts workflows live keeps them right, whether or not we built yours.
Evals, testing and observability for AI systems
An eval suite is the test suite for a system whose behavior is not deterministic: graded cases that state what correct means for your workflow. We build the suite, wire it to run before anything ships, and add tracing so any production answer can be followed back to its inputs, its retrieved context, and the call that produced it.
The suite states what correct means; a change ships only when it is green. The first run you see will be your own.
Agent monitoring and fleet governance
Once agents act on real systems, someone must know what each one did and be able to stop any of them. We install monitoring that records every action with its context, alerts on thresholds you agree, and gives you the controls to pause, roll back, or retire an agent. The default metric set is named, not vague: eval pass rate, exception-queue depth, latency, cost per run and per period, and tool-call errors. Which of them alert, at what values, and to whom, is written into the scope.
Monitoring watches production against thresholds you agree; crossing one raises an alert, not a surprise.
Bespoke integration and MCP, with multi-model backends
The value in production AI sits in the integration: models, tools, and your systems of record working under your access rules. We build that connective layer, including MCP servers for your internal tools where the standard fits, with multi-model backends so a provider change is configuration, not a rebuild.
One MCP layer connects multi-model backends to your internal tools; swapping a provider is configuration, not a rebuild.
Guardrails, red-teaming and AI security
A system that reads untrusted documents and holds tool access can be steered by what it reads. We install input and output guardrails, scope tool access to least privilege, and red-team the system before and after go-live: prompt injection, data exfiltration, and the failure modes specific to your workflow.
Untrusted input and model output each pass a guardrail; tool access stays least-privilege inside the boundary.
Cost control for agents and LLM workloads
LLM spend is a unit-economics question: cost per document, per ticket, per run, not one monthly bill nobody can explain. We meter per workflow, route work to the cheapest model that clears your accuracy bar, and set budgets that alert before they are exceeded. Routing is a measured policy, not a preference: the candidate models are scored on your own held-out set, work goes to the cheapest one that clears the bar, and a miss falls back up the list.
Spend is metered per document and per run, with a budget that alerts before it is exceeded and work routed to the cheapest model that clears your bar.
How the retainer works
The eval suite is the contract of correctness: the agreed statement of what the system must get right, kept current by a named loop rather than by intention. Corrections from the review queue become new cases, the workflow owner adjudicates what correct means when a case is disputed, and the suite is versioned with the system, so a release always names the suite it cleared. Monitoring watches production against named thresholds, and changes, a new model, a new document type, a new integration, pass through an agreed change process: evaluated, gated, then shipped. The gate keeps two of the Seven Gates standing after go-live, the accuracy bar and observability, and every change clears them for as long as the system runs. Each month you get a report of what was measured, what changed, and what it cost. We would rather show you a real one; the first report you see will be yours.
Scroll to view the full schematic
Changes to a live system enter at observe and ship only through the gate.
The workflow review is free and tells you what your case would take.
Price and terms
Retainers start at $2,000 USD per month. The figure for your case is set against the scope of what we run, the number of workflows, how much traffic they carry, and how tight the response expectation is, and it is agreed in writing before the month starts. Per scope, never per seat, so the price does not move when your team grows.
If the build comes first, it is a fixed price from $30,000 USD with staged payments, and the retainer is optional at the end of it: you own what we built either way. Stop the retainer whenever you like. Your system keeps running on your own accounts, and the bills keep coming from your own providers.
Run retainers start at $2,000 USD per month, set against the scope of what we run.
2026-08-19
The workflow review is free and tells you what your case would take.
What we do not sell
If we walked away tomorrow, your system would keep running and your bills would still come from your own providers. That is the test we build to. We do not resell gateways, dashboards, or seats; the service is the judgment and the work, not the plumbing.
Fair questions
Do we have to rebuild?
No. Instrumentation comes first: an eval suite and monitoring around the system as it stands. The measurements then name the points that need fixing; a rebuild is a last resort you would see coming.
Can you work with what we already built?
Yes. A run engagement is built to start from a system someone else built, live or stalled. We instrument what exists first, so every decision about it is made against measurements instead of recollections.
What happens when it breaks?
Monitoring raises the alert, the change process already holds a rehearsed rollback for everything we shipped, and a person acts on it. The coverage window and the response expectation are written into each retainer's scope, like the price: set per scope, not asserted on a website. Every incident then lands in the monthly report with what happened, what was rolled back, and what changed so it does not repeat.
Who owns what?
You do. Everything we build in your tenant, code, eval suites, configuration, documentation, is yours, and the contract says so in writing.
What does it cost?
Retainers start at $2,000 USD per month, set against the scope of what we run: per scope, not per seat. A build is a fixed price from $30,000 USD with staged payments. The workflow review is free and tells you what your own case would take; the figures and their basis are in the price section above.
How does security work?
Delivery work happens in your tenant, under your keys; your documents and systems stay yours. The full posture, including the two kinds of data and how each is handled, is on the data and security page.