Most stalled pilots were set up to stall. What to settle before a pilot starts, how to choose one worth running, and how to run it so it can scale.
Too many AI pilots work as a demonstration and die as an operation. The model does what the supplier said it would. Then nothing happens: no one owns the result, the data it needed was never going to be reliable, and the people who had to change how they work first heard about it in a project update.
The wider numbers point the same way. McKinsey reports that 88% of organisations use AI in at least one function but only about 1% consider themselves fully mature. Gartner's 2026 survey of 140 senior supply chain leaders found that 56% name legacy integration and 50% name limited expertise as major obstacles to scaling AI. This article is deliberately about what to do rather than what the research says. It draws on the method behind C-Insight's AI Bootcamp and Use Case Workshop and AI Readiness Assessment, which were built around the same causes of failure.
The causes are consistent. No clear business objective. A process that was already broken and was automated anyway. Data that does not exist in a usable form. No owner after go-live. Workforce resistance that was never addressed. Not one of the five is a model problem. All five can be settled before a pilot is approved, and most can be settled for the cost of a few well-run conversations.
The pattern in the initiatives that do work is also consistent: narrow scope, high volume, repetitive work, a tolerance for occasional error, and a human check on the output. Where value has not been realised, it is usually broad judgement work, low volume decisions, or anything where being wrong is unacceptable. That is a useful filter before any supplier is called.
In most organisations the first obstacle is that nobody means the same thing by AI. Executives have been asked for a position, teams are using tools nobody approved, and proposals arrive that nobody in the room can assess on merit. A pilot approved in that state is a pilot approved on the strength of the most confident voice.
The fix is to describe capability as work, not technology. AI does five kinds of work: it predicts a number, sorts items into groups, extracts information from documents, drafts content, or recommends a choice among options. Ask your own people to describe their jobs in those terms and the conversation changes. It stops being about products and starts being about decisions, volumes and records. One discipline matters throughout: keep product and supplier names out of the room until the use case has survived scrutiny. Tool questions go on a parked board and get answered at the end, in categories, not names.
Asking a room where it could use AI produces answers shaped by whatever people have recently read or been sold. A better method asks about the work. Run three passes over a process. Where is a decision made repeatedly, on similar information, by an experienced person? Where does information get re-keyed, reformatted, chased or reconciled? Where does a document, form or email arrive and have to be read by a person before anything can happen? In a full workshop the three passes are designed to produce forty to eighty finds, which then cluster into twenty to thirty distinct improvement cases.
Then apply four filters, in order, before anyone discusses how a solution might be built. Frequency and volume: does it happen often enough to be noticed? Data existence: does a record of past instances exist that someone in the room could retrieve? Consequence and check: can the process tolerate an occasional wrong answer, and is there a point where a person would catch it? Ownership: is there a named person who would own it after go-live, and would they want to? Each filter removes items. Frequency is usually the most valuable, because it removes the interesting but rare.
Keep the list of what failed the filters. The not worth doing list is as useful as the shortlist, because it is the evidence that the method is capable of an answer that involves buying nothing.
A pilot can be well chosen and still land in an organisation that cannot absorb it. The AI Readiness Assessment tests that across seven dimensions: strategic clarity and sponsorship, process maturity, data quality and accessibility, technology foundations, governance and risk, people and capability, and change capacity. Technology carries the lightest weight, alongside governance and change capacity, on purpose. Process maturity and data carry the most, at 20% each.
Two features of the method are worth borrowing even if you never commission it. First, Level 3, meaning documented and repeatable, is the practical threshold at which a dimension supports an initiative. Second, averages conceal. An organisation with good strategy, systems and people and no usable data will average well and still fail, so the assessment applies gates: any dimension below 2.0 caps the overall position at Not Ready, and data below 2.5 caps it at Conditional.
The other lesson is about evidence. A self-completed questionnaire measures what the people completing it believe, and senior people tend to see reported numbers rather than the manual work behind them. That is why the assessment rates every score by what it rests on: a document sighted, two independent consistent accounts, or a single uncorroborated assertion, which is capped. The interview that matters most is the one with the analyst who assembles the management reports every week, because that person touches the raw data daily.
You can run this yourself in an hour with the right people in the room. Any question you cannot answer is itself a finding.
| # | Question | Why it decides the outcome | Where it comes from |
|---|---|---|---|
| 1 | What number on the scorecard moves if this works? | A pilot with no target cannot succeed or fail, it can only continue. | Strategic clarity |
| 2 | Who is accountable, with authority to fund and to stop? | Pilots without a named owner drift until the budget cycle ends them. | Strategic clarity |
| 3 | Is the process documented, and does the document match what people do? | Automating a process nobody has written down means automating the version nobody agrees on. | Process maturity |
| 4 | How often does it happen, and at what volume? | Rare events are interesting and rarely worth building for. | Filter 1: frequency |
| 5 | Does a record of past instances exist, and can someone in the room retrieve it? | No usable history means no pilot, whatever a supplier says. | Filter 2 and data |
| 6 | What happens when the system is wrong, and who catches it? | Where an error cannot be recovered, the process is not a pilot candidate. | Filter 3: consequence |
| 7 | Who owns it after go-live, and do they want to? | Work that belongs to everyone belongs to nobody. | Filter 4: ownership |
| 8 | Do you know what AI staff already use, and what company information they may enter? | Unapproved use is the pilot already running, without governance. | Governance and risk |
| 9 | Have the people who do the work helped shape it? | Frontline staff who were told rather than asked will route around it. | People and capability |
| 10 | What else is landing in the same teams this quarter? | A pilot that competes with three other changes for the same people loses. | Change capacity |
Source: C-Insight AI Bootcamp and Use Case Workshop and AI Readiness Assessment frameworks, condensed by the author.
Settle the following in writing before the first line of configuration, because they are impossible to bolt on later. Define success as a number on the scorecard, with a baseline. Name an owner with authority to fund and to stop. Decide in advance what scale and stop look like, so a pilot that misses its target ends rather than lingers. Design the handover between system and person deliberately: where does the human check sit, and what is the exception route? Redesign the process, do not automate the old one. Automating a poor process produces a faster poor process, and most of the available value sits in the redesign.
Treat autonomy as something earned in stages. Start with the system recommending and a person deciding, record how often the person overrides it, and widen the permission only when that record justifies it. And involve the frontline early. The people doing the work spot the exception cases that never made the process document, and they are the ones who will decide in the first month whether the thing is used.
The AI Bootcamp and Use Case Workshop is for the organisation that does not yet know enough to specify what it wants investigated. It builds shared understanding across the leadership group and the people who would own the work, then converts it into a shortlist of improvement opportunities drawn from your own processes, with a not worth doing list. It runs as a full two day format, a compressed one day format, or a half day Executive Briefing for leadership teams that need a common understanding first. The briefing is education only and produces no shortlist.
The AI Readiness Assessment is for the organisation that has initiatives in front of it and is not confident it could land them: a pilot that stalled, a proposal on the table, a board question nobody can answer. It runs over two to three weeks, with eight to sixteen interviews depending on the size of the organisation, and produces a readiness report, a scoring workbook, a prioritised gap list of ten to fifteen items, and an executive summary. There is no free or self-serve version, because a form cannot detect the gap between what executives believe and what the frontline does.
Neither engagement values the opportunities, builds a business case, or recommends any product or supplier, and the assessment is not a cyber security review. The two answer different questions. An organisation can be entirely ready and have nothing worth doing, or have plenty of opportunity and be unable to execute.
For executives and practitioners, the cheapest point to fix a failing pilot is before it starts. Ten questions and three honest conversations cost far less than a six month pilot that ends in a shrug, and the credibility you keep is worth more than the budget.
For advisers, the pilot graveyard is a diagnostic. When a client says the technology disappointed, ask who owned it after go-live, what the baseline was, and whether the process was redesigned. The answers usually point to the organisation, not the tool.
For investors and boards, ask for the not worth doing list and the owner of every live pilot. A company that can show what it declined, and why, has a method. A company that can only show what it is running has a list of experiments.
The industry spends too much effort choosing tools and too little deciding whether it is ready to use one. Readiness is unglamorous. It is documentation, data ownership, a named accountable executive and enough capacity to land one more change, and none of it makes a good slide. But it is the part that decides whether the pilot becomes an operation.
The counter-argument is fair. Moving fast matters, assessment can become an excuse for delay, and some organisations learn more from a cheap, fast experiment than from three weeks of interviews. I agree for low-stakes, contained experiments with no sensitive data and an obvious owner. Run those. But once a pilot touches core processes, customer data or several teams, the cost of discovering the gaps by failing rises sharply, and I would rather an organisation learn that in a workshop than in a write-off. I also have an obvious interest in this view, which is why I have set out the method in enough detail for you to run the ten questions without me.
Weighing a pilot, a stalled trial or a proposal? We will give you an honest read on where to start, including telling you when you do not need us.