From AI pilot to production: what changes when the system has to last
A pilot proves that a model can do the task. Production proves that a company can rely on it every day. Between the two sit integration, validation, ownership and cost, and that is where many AI projects stop.
- By
- Riccardo Benedetti, CTO
- Published
6 min read
In short
- AI pilots rarely stall because of the model. They stall on data access, integration, exceptions, ownership and cost.
- In production the model is one component, surrounded by deterministic validation, human review of exceptions and an audit trail.
- An evaluation set built from real documents is the only reliable way to know whether a change makes the system better.
- Cost per processed document, not cost per demo, decides whether the system can scale.
- Every production AI system needs a named owner on the business side and one on the technical side.
Many companies have run an AI pilot by now. Plenty of them were encouraging: a model read invoices, summarised contracts, answered questions from a manual. Far fewer of those pilots have become part of how the company works.
The gap is rarely the model. It is everything the pilot was allowed to ignore.
Why pilots stall
A pilot is designed to answer one question — can the model do this? — and it is right to cut corners to answer it quickly. The trouble starts when the corners it cut become the project.
- Data access. The pilot worked on a folder of sample files. Production needs the real mailbox, the real archive and the permissions to read them, agreed with IT and security.
- Integration. The pilot produced a spreadsheet or an answer in a chat window. Production needs the result inside the ERP, the CRM or the transport system, in the format they accept.
- Exceptions. The pilot measured the cases that worked. Production has to decide what happens with the ones that don't: the blurred scan, the invoice in a new layout, the email with two orders and a correction.
- Ownership. The pilot belonged to an innovation team. Production needs someone in operations who answers for the results, and someone in IT who answers for the system.
- Evaluation. The pilot was judged on impressions. Production needs measurements that can be repeated after every change.
- Cost at scale. A demo costs cents. Thousands of documents a month, with retries and long attachments, cost real money, and that cost must be known before it appears on an invoice.
What production requires
In a production system the model is one component among many, and usually not the largest. Around it sit the parts that make its output reliable.
Integration with the systems of record
The result of an AI step has to land where the company already keeps its records. That means writing to the ERP or the transport management system through the interfaces they offer — services, staging tables, import files — with the same validation a person's input would receive. The system of record stays the system of record; we describe the pattern in bringing AI into ERP and legacy systems.
Deterministic validation around the model
A model is probabilistic; accounting is not. Every value the AI extracts should pass rules that do not depend on the model: totals that add up, VAT numbers that exist in the master data, dates that make sense, codes that match. When a rule fails, the case stops and waits for a person.
One example from our work. In a supplier-invoice flow for a freight forwarder, the first step is not extraction: it is checking that the PDF really is an invoice. A mailbox that receives invoices also receives reminders, statements and duplicates, and data extracted from them produces errors that look legitimate. Read the case study.
Another. In a pilot for a transport and logistics group, the orders read from customer emails are validated all-or-nothing: if one order in a request fails its checks, none of them reaches the transport management system. A partial import is harder to repair than no import at all. Read the case study.
Human review of exceptions
Human oversight does not mean that a person checks everything; that would remove the benefit. It means the system knows when it is unsure, sends those cases to a person, and shows the source document, the extracted data and the rule that failed. The reviewer's decision is recorded, and the difficult cases become part of the evaluation set.
Over time, the share of cases that go to review should fall as rules and examples improve. If it doesn't, that is information too: the process may have more variety than anyone had written down.
Evaluation sets and monitoring
Before production, build a set of real examples with the correct answers, the difficult ones included. Run it before every release and every change of model or prompt. In production, follow three numbers: the share of cases that go to review, the corrections people make, and the cost per case. A slow drift in any of them is the earliest warning you will get.
Audit trail, identity and permissions
For every case, keep what came in, what the model produced, which rules ran, who approved, and what was written to which system. Put the application behind the company's single sign-on and give it roles, so that an assistant only shows each user what that user is allowed to see. These are the questions auditors will ask, and so do the GDPR and the EU AI Act.
Cost control
Measure the cost per processed document or conversation from the first day of the pilot. Use the smallest model that meets the accuracy target, avoid sending the model what it does not need to read, and keep the option to switch providers. A model-agnostic design is also a cost decision. And count everything: retries, long attachments and the reprocessing that follows a change of model.
A named owner
Every production system needs two names: a business owner, who decides what “correct” means and handles escalations, and a technical owner, who answers for availability, security and change. Without them the system has users, but nobody responsible for it.
A practical checklist
Before we call an AI system ready for production, we check that each of these points has an answer:
- The process, its volumes and the cost of doing it by hand are written down.
- The system reads real data through agreed, permissioned access.
- Results reach the system of record through its own interfaces, and every import is safe to repeat.
- Deterministic rules validate every value that matters, and a failed rule stops the case.
- Uncertain cases go to a person, with the source document and the reason visible.
- An evaluation set of real examples runs before every release and every model change.
- Accuracy, review rate and cost per case are monitored in production.
- Every step is written to an audit trail, and access follows single sign-on and roles.
- Data stays in the agreed region, and every processor is documented.
- There is a business owner and a technical owner, by name.
How long it takes
The honest answer is that it depends less on the model than on the organisation: how quickly data access is granted, how clean the master data is, how many exceptions the process really has. What we can say is where the time goes. The AI step is usually the first part to work. Integration, validation, review screens and monitoring take longer, and they are what makes the system last.
In one sentence
A pilot answers whether the model can do the task; production answers whether the company can rely on it every day.
Where to start
Pick one process with a clear cost and a clear owner — supplier invoices keyed by hand, orders re-typed from emails — and design the production system from the start, even if the first release is small. A pilot built this way does not need to be rebuilt to go live. It only needs to grow. Start with the case whose results you can measure, not with the most impressive one.
This is how we work in enterprise AI, with the architecture and operations that keep a system running once it is live.