AI pilots are easy. Getting them into production is the hard part

Enterprise AI production

AI has become synonymous with eye-watering levels of investment. Gartner expects infrastructure spending to exceed $2.7 trillion this year, with 91% of organisations already planning to increase their AI investment. To put that figure into perspective, it eclipses the entire annual economic output of nations such as Saudi Arabia, the Netherlands, and Switzerland.

Behind the headlines, however, getting from concept to pilot to production is by no means guaranteed. Investment alone offers no promise of operational success. Industry luminaries such as Gartner, McKinsey and Deloitte have all identified a stark “production gap”: whilst roughly 70-80% of enterprise AI pilots get off the ground, only 20-30% ever reach production at scale.

To understand why this gap is widening, we must look at how enterprise AI itself is evolving. We are rapidly moving past the initial waves of conversational Generative AI: chatbots that merely output text, summarise documents, or draft emails. We are entering the era of “Agentic AI”, where, unlike passive language models, agentic systems are built for operational autonomy. Agents reason through complex problems, break tasks into logical sequences, call software tools, and autonomously execute multi-step workflows directly across production environments. Giving AI systems the agency to perform such tasks unlocks incredible productivity, but also introduces a fundamentally new risk profile that legacy architectures were never designed to handle.

While Generative AI unlocks instant knowledge access, Agentic AI introduces operational execution. Together, they represent the future of enterprise productivity, yet both hit a hard wall when moving from isolated testbeds into live production environments. So what’s going wrong? AI pilots are generally designed to operate in a contained setting with limited access to enterprise information and systems, often necessitating the use of dummy or incomplete datasets until the concept has been proven. Moving to production, however, forces these autonomous workloads out into the open. Three friction points tend to trip up the transition from pilot to production: getting the underlying data and systems in order, settling who owns the service once it’s live, and securing data and infrastructure access.

Data and systems

Moving an AI pilot into production requires bridging the gap between contained isolation and live operational reality. This means connecting models and agents directly to existing applications, databases, and core business workflows, a transition that exposes friction points across both Generative and Agentic AI.

For conversational Generative AI, production deployment means connecting models to live corporate knowledge bases. This introduces immediate risks around data lineage, access controls, and context leakage, as organisations must ensure that sensitive customer data or proprietary IP never exits the corporate perimeter. For Agentic AI, the integration challenge escalates significantly because these systems do not merely read data; they act upon it. Connecting autonomous agents to live operational environments requires robust identity boundaries and system integration layers that legacy architectures were simply never designed to support. Many established systems lack the integration capabilities required for automated agents to safely invoke actions, making production integration far more complex than it initially appeared during the pilot phase.

Once AI supports a live process, organisations need to control access and monitor how the service performs. The shift from pilot to production also raises broader ownership and governance questions, particularly when an innovation team built the pilot, but a separate IT function must subsequently operate it.

Infrastructure issues

While practically anyone can spin up a prototype using public AI systems, production systems require an infrastructure environment that reliably scales, governs, and secures workloads as adoption expands. A fundamental misstep when moving beyond the pilot phase is assuming that selecting a more capable frontier model solves the underlying operational challenge. Model alignment and prompt engineering shape what an AI model or agent intends to do, but only the underlying infrastructure controls what it is able to do.

To see how quickly requirements change, consider a typical rollout. An employee could design and run a Generative AI customer-service pilot using a small document set, demonstrating that it can answer routine questions more quickly. Turning that into a production service is a completely different undertaking. The assistant needs secure access to live account data via existing systems, low-latency processing, and predictable compute capacity to scale with customer demand. If that assistant evolves into an autonomous agent capable of modifying account details or initiating refunds on a customer’s behalf, the stakes rise much further. The infrastructure must enforce strict runtime sandboxing so the agent cannot exceed its authorised boundaries. In both cases, the organisation needs clear controls over what information the service can access, with IT teams able to oversee its performance once it becomes part of an everyday business process.

"Ultimately, bridging the production gap is not just about writing better prompts or choosing larger models; it is about grounding enterprise AI in hardened, inspectable infrastructure and clear operational ownership that yields total control."

The list of infrastructure requirements can grow very quickly. Teams need a consistent way to deploy and manage AI across all parts of the estate where it needs to run; spanning developer workstations, core data centres, sovereign regional clouds, and resource-constrained edge environments. Compute capacity, GPU efficiency, and workload placement become practical concerns once AI moves beyond a limited trial. The same environment should also allow organisations to apply security and access policies consistently, ensuring sensitive data and execution runtimes remain under close control.

This is why the production gap cannot be resolved by just choosing a specific AI model alone. It’s not necessarily about whether the organisation in question is using ChatGPT, Claude, or an open-weight local alternative; if the underlying infrastructure cannot accommodate, secure, and govern the workload as part of the wider enterprise estate, a pilot is unlikely to evolve into a durable production service.

Plan for production early

Given the friction many organisations face when transitioning from pilot to production, it’s vital to assess the infrastructure, security, and operational requirements as early as possible.

This begins with identifying the broader operational environment in which the service will reside. A common platform approach can help businesses deploy both Generative and Agentic AI workloads across their existing infrastructure without creating an isolated “AI silo”. This process should allow teams to assess workloads consistently while retaining control over where sensitive data will be processed and, of course, cover what enterprise information the service would need before providing access to live systems.

Open-source foundations make this long-term flexibility far easier to achieve. They grant organisations the architectural freedom to swap individual components, whether that means pivoting between silicon vendors, switching models, or updating software frameworks, as requirements shift, rather than being tied to a single vendor’s roadmap. The same thinking applies to where workloads actually run. Whilst some AI services sit comfortably in the cloud, others need to stay on-premises or within sovereign regional clouds due to latency, cost, or strict data residency reasons, and a growing number will need to process data at the edge, where the data itself is generated. Planning for this flexibility early, including hybrid and multi-cloud options, makes it far easier to move a service into production without a costly re-platforming exercise further down the line.

It’s tempting to focus planning purely on how to mitigate risks, but what happens if a new service is wildly successful?

Central to this ‘nice to have’ challenge is that a service designed for a small test group will encounter materially different compute and operational demand when made available more widely. IT teams need time to determine how compute capacity will scale to meet this demand and how the service will be monitored once live. For Generative AI, this means tracking GPU health and utilisation, response latency, and token throughput; for Agentic AI it means auditing reasoning and enforcing safety and security guardrails so unexpected changes in behaviour are caught before they affect a live business process. Addressing these questions early makes it easier to uncover integration problems before a successful pilot creates pressure for a rapid rollout.

Who owns AI once it enters production?

The practicalities of moving AI into a live environment introduce critical operational considerations, with organisational ownership amongst those with the potential to throw projects off course. A major benefit of modern AI tools is that they lower the barrier to innovation. A pilot may begin with a small agile innovation team, or a single line of business exploring a specific opportunity or the basis of a new idea.

However, deploying that pilot into production may very well create a service that directly impacts customer experience, regulatory compliance, or an established internal process. The organisation must decide who is accountable for the service once the decision has been made to integrate it into day-to-day operations. This requires establishing a clear operational handoff between the innovation teams who prototyped the service and the IT, platform, and SecOps functions responsible for maintaining, observing, and defending it over time.

Ownership must include formal authority to approve material changes to how the service operates. Governance cannot end at launch; it must continue as workloads gain broader access to enterprise systems. For Generative AI, governance focuses on monitoring model accuracy, data lineage, token expenditure, and GPU health. For autonomous Agentic AI, true governance demands deterministic safety rails, enforcing human-in-the-loop approvals before an agent can merge production code, modify database schemas, or execute critical business transactions.

This kind of clear accountability is vital, not only to mitigate operational risk, but to help the organisation continually measure whether the AI service is generating enough business value to justify ongoing investment. Ultimately, bridging the production gap is not just about writing better prompts or choosing larger models; it is about grounding enterprise AI in hardened, inspectable infrastructure and clear operational ownership that yields total control.

Rhys Oxenham, VP and GM of AI, SUSE

Rhys Oxenham

Rhys Oxenham, VP and GM of AI at SUSE, where he leads product strategy and development across the company’s AI portfolio. Before moving into this role, he spent three years overseeing SUSE’s Edge and Telco Engineering teams, delivering mission-critical infrastructure in some of the most demanding operating environments worldwide.

Author

Scroll to Top

SUBSCRIBE

SUBSCRIBE