AI Workloads Leaving the Public Cloud Have Four Things in Common

AI Workloads Leaving the Public Cloud Have Four Things in Common

Key takeaways

  • Enterprise IT decision-makers are pulling AI workloads back from the public cloud to meet strict compliance requirements and avoid consumption-based pricing.
  • The workloads share four traits: steady-state utilization, heavy data movement, latency sensitivity, and regulatory or sovereignty constraints. 
  • Instead of completely abandoning the cloud, organizations are adopting hybrid models that enable them to match workloads to the most appropriate environment.

The trend is clear: Organizations are moving AI workloads out of the cloud. As it turns out, some of these workloads were never a good fit for the cloud model. 



This is not a niche movement. According to a Cloudian survey, 93% of enterprises with AI apps in production have already repatriated AI apps from the cloud, are planning a migration, or are actively considering it.1 These organizations are not giving up on the cloud entirely. They’re moving to a hybrid infrastructure where some workloads will run on dedicated hardware outside of the cloud.

As a growing number of enterprises explore repatriation, it’s important to understand when it makes sense and when it doesn’t. After all, the cloud can continue to provide flexibility and scalability for workloads that need burst capacity, even if other workloads are better hosted elsewhere.

The Problem with Cloud Pricing

The cloud consumption billing model solved a real challenge. Before that approach was introduced, teams had to buy enough resources to accommodate their predicted peak loads—and then continue to pay even when demand was low. With the metered pricing model, those teams no longer paid for idle resources.



The problem is that the model assumes that there will be periods of low demand. But in fact, some workloads—including some production AI workloads—never experience those troughs. Certain AI workloads use compute, memory, storage, and networking resources continuously. Nothing changes from one day to the next. As a result, paying consumption-based rates that were designed for occasional bursts is not cost-effective. 



The same logic applies to AI services priced by the token. Many teams access large language models through managed cloud AI services that charge for every token processed, on both input and output. For a pilot, that metered model makes sense. For a production application handling millions of requests a day, per-token fees add up quickly and are difficult to forecast. Running open-weight models on dedicated bare metal GPU servers replaces those variable charges with a predictable, fixed infrastructure cost. Workloads that still call frontier models for certain tasks will continue to incur some per-token fees, but routing routine requests to self-hosted models can reduce those fees significantly.


Organizations are overpaying for cloud resources across many workloads, not just AI workloads. According to the Flexera 2026 State of the Cloud Report, organizations estimate that 29% of their cloud spending is wasted.2 With the majority of large enterprises now spending more than five million dollars each month on cloud services, that waste is significant.3

Losing Control of Workloads in the Cloud

Cost can certainly be a key reason for repatriation. Yet data sovereignty is often the deciding factor. Organizations want greater control over their data than public clouds can offer. They need to enhance security and maintain compliance with strict data privacy regulations. Cloudian found that 91% of organizations deploying AI apps that touch sensitive company data would choose on-premises, private cloud, or hybrid infrastructure over public cloud environments.4

Which Workloads Are Moving

Repatriated AI workloads typically share four common traits. Assessing each workload according to these traits can help teams solidify repatriation decisions.

Steady-state utilization

Public cloud pricing is geared for workloads that occasionally burst. But AI inference workloads, for example, typically do not. These workloads run continuously, at high utilization. Demand does not suddenly spike. Any similar steady-state workloads—whether they involve AI or not—are a better fit for dedicated hardware than for cloud environments. When those inference workloads rely on a managed model service, always-on demand also means a constant stream of per-token charges.

High-volume data movement

AI workloads are data-intensive. The data that is part of training sets, feature stores, embeddings, model checkpoints, and inference logs is frequently moving—and cloud providers charge for that movement. Repatriating AI workloads that move a lot of data can reduce or eliminate data egress fees.

Latency sensitivity

Many AI apps depend on near-real-time responses to input. When organizations keep workloads in centralized cloud data centers, data might need to travel longer distances to and from users, creating unacceptable latency. And when those workloads must compete for cloud resources with other tenants in shared infrastructure, latency can increase. Real-time inference, fraud scoring, industrial control, and interactive AI apps would all benefit from dedicated infrastructure, located close to users.

Sovereignty and sensitivity

For some organizations, data sovereignty and residency laws mandate where AI workloads and data must reside. Complying with national and regional laws can be difficult when running those workloads in the cloud. Similarly, strict data security and privacy laws might cause some organizations to reconsider cloud environments. Keeping workloads in environments that they directly control can provide greater assurance of protection and a clearer path to compliance.

Implementing a Hybrid Model 

While many organizations are moving AI workloads back from public cloud environments, they are not completely abandoning the cloud. Teams recognize that some AI workloads do in fact burst—and therefore benefit from the elastic resources and flexible, consumption-based pricing offered by cloud providers. For example, an airline’s AI-powered customer service chatbot might need to scale up instantly if poor weather causes a cascade of flight delays.

To accommodate both steady-state and “bursty” workloads, organizations are readily adopting hybrid models. In fact, Gartner predicts that more than 40% of leading enterprises will have adopted hybrid architectures by 2028.5 Instead of leaving the public cloud, organizations are simply becoming more strategic in terms of where they run specific workloads.

Where Do the Workloads Go?

There’s no doubt that organizations are moving some AI workloads out of the cloud. But where are those workloads going? Not all are moving to on-premises data centers, which organizations might need to expand to accommodate repatriation plans. Some might be shifted to a colocation facility, for example, where an organization can run dedicated hardware in someone else’s data center on a monthly contract.

The decision to move a workload to one location or the other is critical—and it could determine whether a repatriation project succeeds. Part two of this blog post will explore that essential repatriation decision.

FAQs

Q: What is cloud repatriation? 
A: Cloud repatriation involves moving workloads out of a public cloud environment and onto infrastructure that an organization controls directly. That non-cloud environment could be on premises, in a colocation facility, or another facility that hosts dedicated hardware. Organizations often repatriate workloads to satisfy compliance requirements or to avoid cloud pricing models. 

Q: Are enterprises completely abandoning the cloud? 
A: No. Many are conducting workload-by-workload assessments and repatriating only those workloads that are a better fit for dedicated hardware. With a hybrid model, they can benefit from cloud flexibility for some workloads and bare metal for others.

Q: Which workloads are the best repatriation candidates? 
A: Workloads with steady utilization, high-volume data movement, latency constraints, or data residency requirements are often the best choices for repatriation. AI inference workloads often meet all four criteria.

Q: Is repatriation only relevant to large enterprises? 
A: No. Any organization with always-on workloads and significant data movement, for example, can benefit from shifting those workloads from the cloud. These workloads are not a good match for a consumption-based cloud pricing model. 

Q: Does moving AI workloads to bare metal eliminate token charges? 
A: It can. When you run open-weight models on your own GPU servers, you pay for infrastructure rather than for each token processed. If your application still sends some requests to frontier models through a provider’s API, those calls will carry per-token fees, but self-hosting routine inference can reduce the total substantially.

 

Citations
1Cloudian, Enterprise Survey Finds 93% Are Repatriating AI Workloads or Evaluating a Move Away from Public Cloud, March 2026
2Flexera, 2026 State of the Cloud Report, March 2026
3Ibid.
4Cloudian, Enterprise Survey, March 2026
5Gartner, Gartner Identifies the Top Strategic Technology Trends for 2026, October 2025

Come see what the Hivelocity difference
 means for your organization
Get expert guidance on choosing the right cloud solution for your enterprise needs.
Disaster Recovery
How to Survive When Ransomware Strikes
Don’t Miss What’s Next!
Register for live webinars, join expert AMAs, explore in-person meetups, and more.