Why your AI pilot never shipped
The gap between a demo that works and a system that runs is not the model. It is the four things nobody scoped.
A pilot is judged on whether it can work once. Production is judged on whether it can work every time, unattended, against messy real data, without doing something you would have to apologise for. Those are different problems, and the second one is where most pilots quietly die.
The four things nobody scoped
- The unhappy path: what the agent does when it is unsure, when the data is wrong, or when the tool it needs is down.
- The handoff: the exact point a human takes over, and how they are told.
- The ownership: who holds the keys, the prompts, and the data when the consultant leaves.
- The proof: the one metric the system exists to move, instrumented from day one.
If a demo cannot name the metric it moves and the moment a human steps in, it is not a pilot. It is a screenshot.
The fix is to scope the boring parts first. We commission a first system as a fixed piece of work with those four answered up front, so the demo and the deployment are the same thing.
Want this graded for your own stack?
A systems audit runs your operation against exactly these dimensions and hands you the report.
Request a systems audit