AI and taking it into use · 9 min·
Nobody owns the stretch between pilot and production
We have taken part in three surveys of how AI is used in the Finnish health and social care system. During them we have interviewed dozens of researchers, clinicians, data management specialists and executives. Among other things, we have asked what has happened to the models they developed.
We expected to hear about technical problems. They were barely mentioned.
University hospitals are developing good models for imaging, automated documentation and clinical prediction. They work. Research is being done at an accelerating pace, and the expertise is there. What is missing is a person whose job description includes taking a finished model from research into use.
The same observation came up so many times that we stopped recording it separately.
A figure circulates in public according to which around 95 per cent of generative AI pilots never reach production.¹ It can be read as proof that the technology has been oversold. Our material says something else. Pilots do what they are asked to do, and then they end.
A pilot does exactly what it should
The job of a pilot is to show that a solution can work in a limited setting. Production readiness is not part of that job. A research group is not responsible for continuous use, for maintenance or for regulatory compliance, and it should not be.
The fault lies in what happens next.
In the interviews the pattern repeated almost without exception. The pilot ends. The project ends. The solution waits for a next stage whose criteria have not been written and for which no owner has been named. After a while the next project starts down the same path from slightly different premises.
The result is duplicated work and lessons lost between organisations. We have seen several times how similar solutions are developed in parallel in different regions. The most surprising part has been how surprised people are when they are told about it.
Starting new pilots before the onward path is defined adds to the fragmentation.
Why the benefit disappears on the way
Benefits measured at task level are real. In the customer support of one large company, problem resolution became 14 per cent faster on average, and productivity among inexperienced and lower-performing staff improved by 34 per cent.² In software development, teams that use AI heavily merge 98 per cent more changes than others.³
At organisation level the effect is not visible.
The same study reports review times in those teams being 91 per cent longer.³ Speeding up one step moved the bottleneck to the next one. At macro level, some estimates predict total factor productivity growth of under one per cent over the next ten years.⁴
This is not new. Robert Solow noted in 1987 that the computer age was visible everywhere except in the productivity figures.⁵ General purpose technologies such as electricity and the computer have historically taken decades before their effect showed at the level of the national economy.⁶ The same phenomenon is known as the productivity J-curve: a new technology requires investment in skills, ways of working and processes, and in the early phase productivity can even fall before it rises.⁷
Boston Consulting Group proposes a 10–20–70 ratio for successfully taking a solution into use. Ten per cent on algorithms, twenty on technology and data, seventy per cent on changing processes and people’s work.⁸ Most of the organisations whose work we have seen from the inside spend their resources the other way round.
What taking a model into production requires
Production use is not a single obstacle. Regulatory expertise, data protection expertise, a production environment, security, lifecycle management and a procurement and contract model that suits this kind of work all have to be in place at the same time.
If even one is missing, the solution stays an experiment. A single hospital or wellbeing services county does not typically have the resources, the mandate or the structure to keep all of these standing at once. Few suppliers have them all together either.
Some of the requirements are not voluntary. Under Article 4 of the EU AI Act, providers and deployers of AI systems “shall take measures to support the development of AI literacy of their staff”. The obligation was relaxed in July 2026: “this obligation does not require providers or deployers to guarantee any specific level of AI literacy”. Measures are still required. In the field, the obligation is in our experience poorly known. The minimisation requirements of the GDPR and of the Finnish Act on the Secondary Use of Health and Social Data, and in future the EHDS, sit in direct tension with the wish in data-driven research for broad datasets that reveal new connections. That tension is not resolved by demanding more of the other party. It is resolved when a research question can be translated into a concrete data requirement, and that is a rarer skill than one would think.
Security is a layer of its own, and it has changed. Language models bring a new attack surface: natural language. We have come across bots that were in production and open to the public internet, and whose weaknesses come straight off OWASP’s top 10 list for language model applications: prompt injection, system prompt leakage, unbounded consumption.⁹ When the same bot is also given access to the organisation’s data and tools, the risk class is different. It is worth assuming that every bot will sooner or later be the target of a prompt injection, and limiting its permissions from that assumption.
What is needed instead
A repeatable structure. These work together, not as separate measures.
A safe experimentation environment where models can be tested with real data without production impact. And an agreed gate from there into production. Without the gate, the environment itself turns into a permanent experiment.
A model registry: which models are in use and in development, versions, owners, purpose, validation status, risk class, production status. A registry starts to work when it is used for prioritisation and approval, not when it has been filled in.
A reference architecture: how AI is connected to the environment. Data flows, access rights, logging, monitoring, MLOps. Without this, every project rebuilds the same foundation in a slightly different way.
Approval gates and a decision rhythm. Go/no-go criteria are agreed in advance, and at regular intervals a decision is made to continue, to redirect or to stop. Missing stop criteria cost as much as missing criteria for going ahead, because resources stay tied up in projects whose conditions are not met.
Lifecycle operation. This is the one most often forgotten. Failure typically happens in the maintenance phase, not in development. Model drift, updates, version control, service level and continuous monitoring are properties of a production service, not the last task of a project.
On top of these, content ownership is needed, and the best example of that is intelligent search on an intranet. A demo is technically easy and quick to build. Underneath, it is a data quality and content discipline project. If the guidance is contradictory, out of date or without metadata, the model returns uncertain answers, and the user is left with the feeling that the AI is making things up. It is only reflecting the confusion of the source material. A search model does not produce one truth from sources that hold several.
Measure the work that moved out of sight
One thing deserves a separate mention, because it is cheap to do and because it changes the nature of the assessment.
A pilot can look as if it works when the work has only moved somewhere it is not counted: into manual checks, workarounds and patching done by people. Measuring workarounds, manual work and operational friction therefore belongs in the assessment of a pilot in the same way as accuracy figures. Without it, you scale a solution whose real costs only become clear at the hundredth user.
And the definition of success has to be verified use. The project has an owner, the solution is in use, and the benefit is tracked. If those three are not met, the project has produced knowledge.
The same observation elsewhere
The systematic literature review by Abdelwanis et al. (2026) brought together 92 studies and identified 16 barriers to taking AI into use in healthcare. The barriers are organised into a Human–Organization–Technology framework, in which technology is one of the three groups.¹⁰ In our material the emphasis is clear: the central problems are structural and organisational.
That is good news, because structures can be decided. Criteria, responsibilities and ownership can be agreed before the next pilot.
There is one question we cannot answer. We do not know whom the missing role belongs to. The purchaser, the supplier or a national actor. All three have been tried somewhere, and none of the models has yet proved clearly better. This much can be said. As long as the role belongs to everyone and therefore to no one, every new pilot is an investment whose return depends on whether someone happens to be interested enough to do it alongside their own job.
Sources
- 1.MIT NANDA (2025). *The GenAI Divide: State of AI in Business 2025.*
- 2.Brynjolfsson, E., Li, D. & Raymond, L. R. (2023). Generative AI at Work. NBER Working Paper 31161. Published version: *The Quarterly Journal of Economics* 140(2), 889–942 (2025). doi:10.1093/qje/qjae044.
- 3.Faros AI (2025). *The AI Productivity Paradox Report 2025.* Published 23 July 2025.
- 4.Acemoglu, D. (2025). The simple macroeconomics of AI. *Economic Policy* 40(121), 13–58.
- 5.Solow, R. (1987). ”We'd better watch out.” *New York Times Book Review*, 12 July 1987.
- 6.David, P. A. (1990). The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox. *American Economic Review* 80(2), 355–361.
- 7.Brynjolfsson, E., Rock, D. & Syverson, C. (2021). The Productivity J-Curve: How Intangibles Complement General Purpose Technologies. *American Economic Journal: Macroeconomics* 13(1), 333–372.
- 8.de Bellefonds, N., Grebe, M., Luther, A. et al. (2024). *Where's the Value in AI?* Boston Consulting Group, 24 October 2024.
- 9.OWASP Foundation. *OWASP Top 10 for LLM Applications 2025.* Liikenne- ja viestintävirasto Traficom & Huoltovarmuuskeskus (2026). *Tekoälyagenttien kyberturvallisuus: ohjeistus tekoälyagenttien turvalliseen suunnitteluun ja toteutukseen.* Traficomin julkaisuja 5/2026. ISBN 978-952-425-008-5.
- 10.Abdelwanis, M., Simsekler, M. C. E., Gabor, A. F., Sleptchenko, A. & Omar, M. (2026). Artificial intelligence adoption challenges from healthcare providers' perspectives: A comprehensive review and future directions. *Safety Science* 193, 107028. doi:10.1016/j.ssci.2025.107028.