Designing a pilot that can fail

Most AI pilots succeed. That is the problem with them. A pilot that cannot fail is not an experiment — it is a purchase that has already been decided, with a period of theatre in front of it.

A real pilot has a hypothesis, a baseline, a fixed end, and a written condition under which you stop. All four, or you are not going to learn anything.

The four parts of a pilot that can actually fail
The four parts of a pilot that can actually fail

Write a hypothesis, not an intention

"We will pilot AI for claim summarisation" is an intention. It cannot be wrong.

"Using AI-generated claim summaries will reduce average review time from 18 minutes to under 12 minutes, without increasing the rate of review errors found in QA sampling" is a hypothesis. It names a metric, a direction, a target, and a guardrail — and it can turn out to be false, which is the entire point.

The guardrail clause is the part people leave out and the part that catches the real failures. Almost any process gets faster if quality is allowed to fall, and a pilot measuring only speed will report success in exactly the case you would most want to know about.

Measure the baseline before you start

You cannot demonstrate an improvement against a number you estimated afterwards, and "it feels faster" is what most AI deployments have instead of evidence.

Kavita spent the first two weeks measuring: average review time per claim file, from a sample of 60. QA error rate, from the existing monthly sample. Reviewer-reported difficulty, on a one-to-five scale, because it turned out to matter later.

Two weeks of measuring nothing new. It is the least popular part of any pilot and the only reason the result means anything. If you take one thing from this lesson: the baseline is not optional, and it must be measured before anyone touches the tool.

Fixed window, fixed scope, one team

Six to eight weeks. Long enough to get past the novelty — the first two weeks of any new tool are unrepresentative in both directions — and short enough that it cannot become permanent by default.

One team, or one sub-process. A pilot spread across three teams gives you three confounded half-results.

The same people throughout. And, importantly, not only the enthusiasts, which lesson 7 explains.

Write the kill criteria down, in advance

This is the discipline that separates a pilot from a procurement, and it takes ten minutes.

We stop if: review time does not drop by at least 20%; or QA errors rise at all; or more than 10% of summaries need to be discarded; or reviewers report they are checking the source file in full anyway.

That last one is the most valuable criterion in the list and it is specific to AI. If the reviewer re-reads the whole file to trust the summary, the summary has added a step rather than removed one. That failure is invisible in a time metric averaged over a small sample, and it is the most common way these deployments quietly fail to help.

Write the criteria before the pilot, circulate them, and get the sponsor to acknowledge them. The reason is human rather than analytical: after six weeks of effort, everyone involved — including you — will want it to have worked. Criteria written in advance are the only defence against that, and they are what let you stop something without it being anyone's failure.

Give it a fair run

Two things sabotage pilots by accident.

Under-training. Handing people a tool with a ten-minute demo and measuring what happens is measuring confusion. Budget real training and let the first week be excluded from the numbers.

Leaving the process unchanged. If the tool produces a summary but the official process still requires reading the full file first, you will measure nothing. The process change is the intervention; the tool is only the enabler.

What Kavita's pilot found

Eight weeks, twelve reviewers, claim summarisation.

Review time fell from 18 minutes to 11.5. QA errors were flat. So far, a clear pass.

The thing she would not have found without the fourth kill criterion: four of the twelve reviewers were reading the full file anyway. Asked why, they said the summaries were good but they had been burned twice by a summary that merged a prior related claim — the exact failure the vendor demo had shown her in lesson 4.

So the deployment shipped with a change: any file referencing a prior claim is flagged and the summary is labelled verify against source. Those are about 15% of files. The other 85% get the full time saving and are trusted.

That is a better outcome than a clean pass would have been, and it exists because the pilot was allowed to surface a problem instead of being run to produce a yes.

Do this today: write the hypothesis and the four kill criteria for whatever you are considering. If you cannot write a criterion that would make you stop, you are not planning a pilot.

← Previous