Why pilots stall
Plenty of UK businesses now use AI in some form. In its July 2026 article on AI in UK businesses, the Office for National Statistics (ONS) reports that self-reported AI use among firms with 10 or more employees rose from around 12% in late 2023 to around 35% by June 2026. The same article names "difficulty identifying business use cases" as one of the most commonly reported reasons businesses delay or avoid AI. Choosing the problem is itself a barrier.
Using AI is one thing; getting value from it is another. McKinsey's global survey, The state of AI in 2025 (November 2025), found that nearly two-thirds of respondents said their organisations had not yet begun scaling AI across the enterprise. Only 39% attributed any EBIT impact to AI, and most of those put it at under 5% of EBIT. In July 2024, Gartner predicted that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025. It cited poor data quality, inadequate risk controls, escalating costs and unclear business value.
A 2025 report from MIT's NANDA initiative, The GenAI Divide, drew attention for its finding that about 95% of the enterprise generative AI pilots it looked at showed no measurable return. Its method (interviews, a survey and a review of public deployments) has been widely debated, and its sample leans towards large enterprises. Treat the figure as a warning sign, not a precise failure rate. Its explanation is still useful: the pilots that failed were mostly brittle tools that did not fit how work was actually done.
Taken together, these sources point to the same causes. Pilots stall when nobody defined what success means, when the data was not ready, when the tool sat outside the real workflow, or when the cost of running it was never weighed against the benefit. All of these can be checked before you start.
Key point
Most stalled AI pilots fail on the problem they chose, the data they assumed and the workflow they ignored. Those are decisions you make in the first fortnight, so that is where to spend your effort.
What a good first use case looks like
Your first project has two jobs. It should deliver some value, and it should teach your organisation how to run AI work. That favours modest, repeatable tasks over ambitious transformations. Look for four properties.
High volume
The task happens often: hundreds or thousands of times a month. Volume turns small time savings into a meaningful total. It also gives you enough examples to test against and enough usage to learn from quickly. Examples include triaging inbound emails, extracting fields from invoices or delivery notes, drafting first responses to routine customer queries, or summarising case notes.
A clear success measure
You can say in one sentence what "better" means and measure it today. "Cut average handling time for tier-one queries from eight minutes to five" works. "Improve customer experience" does not. If you cannot measure the current process, you cannot prove the new one is better.
Tolerable error
AI systems make mistakes, so pick a task where a mistake is cheap, easy to spot and easy to correct. A first draft that a person reviews before sending is a good fit. An automated decision about someone's credit, health or employment is not. Those carry legal and ethical obligations that a first project should not have to carry.
Available data
The inputs the system needs already exist in a digital form you can access, and you have real past examples of the task done well. If the first step is "we would need to start collecting that", choose a different first project.
Scoring: value × feasibility × risk
Once you have a long list of ideas (eight to fifteen is plenty), score each one from 1 to 5 on three dimensions and multiply the scores. Multiplying, rather than adding, means a weak score on any one dimension pulls the total down hard. That is what you want: a valuable idea with no usable data is not a good first project.
- Value (1–5): hours saved, revenue protected, errors avoided or service improved, at realistic volumes.
- Feasibility (1–5): data available, task well defined, fits existing systems, achievable with current tools.
- Risk (1–5, scored so that 5 means lowest risk): errors are cheap and caught, personal data exposure is limited, and few regulatory or reputational concerns apply.
The table below uses illustrative examples to show how the scoring separates ideas. Your own scores will depend on your data and processes.
| Candidate use case (illustrative) | Value | Feasibility | Risk (5 = low) | Score |
|---|---|---|---|---|
| Extract fields from supplier invoices into the finance system | 4 | 4 | 4 | 64 |
| Draft replies to routine customer emails for agent review | 4 | 4 | 4 | 64 |
| Search internal policies and procedures for staff questions | 3 | 4 | 4 | 48 |
| Forecast stock demand across all product lines | 5 | 2 | 3 | 30 |
| Fully automated chatbot answering customers with no human review | 4 | 3 | 2 | 24 |
| Automated screening of job applicants | 3 | 3 | 1 | 9 |
As a rough guide, anything scoring above about 45 is worth taking into discovery. Anything below about 25 should wait until you have more experience. Agree the scores as a group: someone from operations, someone who owns the data and someone who will use the tool. One person scoring alone tends to overrate their own idea.
Data readiness
Poor data quality appears in Gartner's list of reasons projects are abandoned, and it is the issue most often underestimated. Before you commit, answer five questions about the use case you shortlisted:
- Where does the data live? Name the systems, folders or inboxes, and who controls access.
- Can you get it out? An export, an API or a database query, not screenshots or copy and paste.
- Is it representative? Does it cover the awkward cases (handwritten notes, unusual formats, edge-case customers) and not just the tidy ones?
- Do you have good examples? You need at least 50 to 200 past cases with known good outcomes to build a test set.
- Can you lawfully use it? If it contains personal data, check your lawful basis, your privacy notices and whether you need a data protection impact assessment under UK GDPR. The ICO publishes guidance on AI and data protection.
If two or more of these answers are uncertain, either fix them first or pick a use case whose data is already in order.
People and change
The MIT report's explanation for failure (tools that did not fit how work was done) is a people problem as much as a technical one. A tool that works in a demo but adds a step to someone's day will be quietly ignored.
- Name one business owner. This is a person who runs the process today, not someone from IT. They decide what "good" looks like and sign off on go-live.
- Involve the people who do the work from week one. They know the edge cases, and they are the ones who will either use the tool or work around it.
- Be straight about what changes. Say plainly whether the aim is to free up time for other work, absorb growth without hiring, or something else. Uncertainty breeds resistance.
- Put the tool where the work already happens. Inside the inbox, the CRM or the finance system is better than a separate tab.
- Plan the training. Short, task-specific sessions work better than general AI awareness courses. The ONS reports that lack of expertise is a common barrier, so budget time for it.
Key point
If the people who do the work today were not involved in choosing and testing the tool, assume they will not use it. Adoption is part of the project, not something that happens after it.
Build or buy
For many first projects, the right answer is to buy or configure something rather than build from scratch. As Fortune reported, the MIT NANDA research found that tools bought from specialist vendors or delivered through partnerships succeeded about twice as often as those built entirely in-house. The same caveats about the study apply, but the direction matches common sense: someone else has already solved the general problem.
| Option | Best when | Watch out for |
|---|---|---|
| Buy an off-the-shelf product | The task is common across businesses (meeting notes, invoice capture, helpdesk drafting) | Lock-in, data residency, per-seat costs as usage grows |
| Configure a platform or hosted model | The task is common but needs your documents, rules or tone | Quality depends on your prompts, data and testing |
| Build a custom solution | The task is specific to you, sensitive, or a source of competitive advantage | Longer timelines, ongoing maintenance, need for in-house or partner skills |
Whichever route you take, keep ownership of your test set and success measures. They let you compare vendors fairly and switch later if you need to. For the technical choices within a build, see our guide to choosing between prompting, RAG, fine-tuning and custom models.
Scoping a 2–4 week discovery
Discovery is a short, fixed piece of work that ends in a clear go or no-go decision. It is not a study that runs on until enthusiasm fades. A typical shape:
- Week 1: run workshops with process owners and users, build the long list, and map the current process for the top three candidates, including volumes and timings.
- Week 2: score the candidates, check data readiness on the leading one or two, and pull a sample of real data.
- Weeks 3–4 (if needed): run a quick technical spike on real examples to see whether current tools can do the task at all. Draft the success measures, the test set and a cost estimate.
The outputs should be a one-page problem statement, a baseline measurement of today's process, an agreed success threshold, a test set of real examples, a build or buy recommendation, and a costed plan for the prototype. Ending discovery with "not yet" is a legitimate result. It is far cheaper than finding out in month four.
Scoping a 6–8 week prototype
The prototype should be used by a small group of real users on real work. A demo that only runs on hand-picked examples does not count.
- Weeks 1–2: build the simplest version that could pass the test set, connect it to the real data source, and set up logging.
- Weeks 3–4: improve the tool against the test set, add the human review step, and hand it to three to ten pilot users.
- Weeks 5–6: run it on live work alongside the existing process, gather feedback, and fix the most common failures.
- Weeks 7–8 (if needed): measure results against the baseline, estimate running costs at full volume, and write up the decision to scale, change or stop.
Keep the scope tight. Each new feature request during the prototype goes on a list for later, unless it is needed to meet the success threshold.
Key point
Fix the time and let the scope flex. A narrow tool that real users rely on after eight weeks is worth more than a broad one that is still "nearly ready".
What to measure
Measure the same things before and after, using the baseline from discovery. A useful set:
- Quality: accuracy or acceptance rate on the test set, and the share of outputs users accept without major edits.
- Efficiency: time per task, throughput, or backlog size.
- Adoption: the share of eligible tasks where the tool was actually used, and whether usage holds up after the first fortnight.
- Errors and risk: the number and severity of mistakes that reached a customer or a system of record.
- Cost: licence, usage or hosting costs per task, plus the staff time spent reviewing outputs.
Gartner lists escalating costs and unclear business value among its reasons for abandonment. Tracking cost per task from the start addresses both.
Common failure modes
- Starting with the technology. "We should use an AI agent" is not a problem statement.
- No baseline. Without today's numbers, nobody can show the tool made a difference.
- Demo data only. Clean examples hide the messy cases that make up much of the real work.
- No owner. A project run by IT on behalf of a department that never asked for it rarely gets adopted.
- Picking the hardest problem first. High-stakes, low-tolerance tasks belong later, once you have experience and controls in place.
- Endless pilots. No end date and no decision criteria means the pilot never ends and never scales.
- Ignoring running costs. A tool that is cheap for ten users can be expensive at a thousand.
- Governance as an afterthought. Data protection, security review and human oversight are far easier to design in than to bolt on.
Checklist
Before you commit budget to a prototype, check that you can answer yes to each of these.
- The task happens often enough that small savings add up.
- We have a one-sentence success measure and a baseline for today's process.
- A mistake by the tool is cheap, visible and easy to correct, and a person reviews outputs where it matters.
- The data exists, we can extract it, and it includes the awkward cases.
- We have at least 50 real examples with known good outcomes to test against.
- We have checked our lawful basis for any personal data and considered whether a DPIA is needed.
- The candidate scored well on value, feasibility and risk, and not just on one of them.
- A named business owner from the team that runs the process has signed up.
- The people who will use the tool are involved in testing it.
- We have considered buying or configuring before building.
- The prototype has a fixed end date and agreed criteria for scaling, changing or stopping.
- We know what it will cost to run per task at full volume.
The best first AI project is rarely the most exciting one on the list. It is the one that is frequent, measurable, forgiving of error and backed by data you already have, and that is why it ships.
Sources
- Artificial intelligence in UK businesses: 2023 to 2026 (20 July 2026) — Office for National Statistics
- Business insights and impact on the UK economy: 2 July 2026 — Office for National Statistics
- The state of AI in 2025: Agents, innovation, and transformation (November 2025) — McKinsey & Company
- Gartner predicts 30% of generative AI projects will be abandoned after proof of concept by end of 2025 (29 July 2024) — Gartner
- MIT report: 95% of generative AI pilots at companies are failing (18 August 2025) — Fortune
- Guidance on AI and data protection — Information Commissioner's Office