Most public-sector AI pilots do not fail. They simply never end. The model works, the demo lands, everyone agrees it is promising — and then eighteen months later it is still a pilot.
The pilot is not the problem
The technical work in a well-scoped pilot is rarely the hard part. Retrieval over a document corpus, a triage classifier, a drafting assistant — these are solved shapes. What stops them from becoming systems is almost never the model.
Three things stall them, in roughly this order.
1. Nobody owns the outcome
A pilot is usually sponsored by an innovation function and run against a business process owned by somebody else. The sponsor can fund a proof of concept but cannot change the approval workflow it would need to plug into. The process owner did not ask for it and carries all the risk of adopting it.
The fix is unglamorous: name a single accountable owner in the operational line, not the innovation function, before any build starts. If no such person will put their name to it, that is useful information — and it is cheaper to learn it in week one.
2. Success was never defined in numbers
Ask what the pilot was meant to achieve and you often get a capability description rather than a target: "demonstrate that AI can summarise incoming correspondence." That is a statement about the model, not about the organisation.
A usable target looks like: median turnaround on incoming correspondence drops from six days to two, measured on the same queue, over the same case mix, within one quarter. Now the pilot can pass or fail, and passing means something a director can take to a board.
If the pilot cannot fail, it also cannot succeed. It can only continue.
3. Governance arrives last instead of first
Data residency, auditability, model selection and record-keeping obligations get raised at the point of scaling, which is the most expensive possible moment. Public-sector scrutiny is entirely predictable — it should be a design input in the first workshop, not a discovery in month nine.
What to do instead
- Agree the KPI before the architecture. One primary metric, one baseline, one measurement window, all written down.
- Name an operational owner. Someone whose own numbers improve when the system works.
- Bring compliance in at the design stage. Residency and auditability constrain model choice — better to know on day one.
- Scope for production, build the smallest slice. Design the real system, then deliver a narrow but complete vertical of it that runs in the live workflow.
- Set an explicit decision date. On that date the pilot scales, changes, or stops. "Continue as a pilot" is not on the list.
The uncomfortable version
Some processes should not be automated, and some organisations are not ready to change the workflow that would make the automation useful. A discovery engagement that concludes "not this, not yet" has saved considerably more than it cost.
That is a better outcome than a fourth pilot.