Insights · Guide I

The production gap

From pilot to production: five rungs, and the gate at the top of each.

What production means

A system is in production when it sits in the path of real work, has a named owner, is measured against a baseline, is supported when it fails, and is on the register. Anything short of that is a pilot, however many people use it.

The distinction matters because the evidence on pilots is consistent. MIT’s 2025 study found that 60 per cent of organisations evaluated enterprise AI tools, custom or vendor-sold, 20 per cent reached a pilot and 5 per cent reached production: on our arithmetic, one in twelve of those that looked.2 Gartner predicted that at least 30 per cent of generative AI projects would be abandoned after proof of concept by the end of 2025.1 McKinsey’s 2026 survey has 44 per cent of respondents reporting AI scaling across the enterprise, but only 37 per cent attributing any EBIT impact to it.3

Figure 1 The production ladder
  1. Frame

    Is there a problem with an owner and a number?

    A named process owner, a baseline measure, and a written decision on what result would justify going further.

  2. Prove

    Does it work on our data, not a demonstration’s?

    A result on a representative sample of live data, scored against the baseline by the people who do the work.

  3. Pilot

    Does it work in the path of the work?

    Real users at real volume for a fixed period; errors caught and counted; the fallback tested.

  4. Production

    Could it run on a Monday with its builder on holiday?

    Integrated with the systems of record, security reviewed, on the register, supported by a runbook, its owner trained.

  5. Operate

    Is it still earning its place?

    The measure reviewed monthly, drift and incidents tracked, a date for the next review, and a way to switch it off.

Do not start a rung whose gate you cannot describe. Most failed pilots skipped the first. Source: Vardonne

Why pilots stop

Gartner’s reasons are poor data quality, inadequate risk controls, escalating costs and unclear business value.1 Each shows up at a particular point, and each is cheapest to deal with earlier than it appears.

The four causes, where they surface, and where to deal with them
CauseWhere it surfacesDeal with it at
Data qualityThe pilot ran on a cleaned extract; production has to use the live system.Rung 2: prove on live data.
Risk controlsNobody decided who reviews outputs, or the data protection review came last.Rungs 1 and 2: register it, start the DPIA.
Escalating costUsage pricing at full volume, integration and support were never costed.Rung 1: cost production, not the pilot.
Unclear valueThere was no baseline, so no one can say what changed.Rung 1: measure before anything is built.

Exit criteria, written first

Before a pilot starts, write down five things: the measure and the threshold that counts as success; the volume it must handle; the error rate that is tolerable and who catches the errors; who decides; and the date on which they decide. Then hold the date. A pilot that is extended because nobody wants to decide is the most common way good work dies.

What the cases in production share

Across the use-case library, the entries that publish a measured outcome have the same shape: AI sits inside one process, the process has an owner, and the outcome was measured against how things were before. Yorkshire Water’s sewer monitoring and Aviva’s underwriting summaries are different technologies with the same discipline.

Sources

  1. Gartner predicts 30% of generative AI projects will be abandoned after proof of concept by end of 2025. Gartner, 29 July 2024.
  2. The GenAI Divide: State of AI in Business 2025. MIT NANDA, July 2025.
  3. The state of AI in 2026. McKinsey & Company, 25 August 2026.