Your Egress Bill Is Being Written Before the CDN Ever Sees the Traffic

Key takeways

  • Streaming platforms often focus on optimizing CDN contracts to reduce egress costs, but they should instead focus on the origin architecture. 
  • Using hyperscalers for the origin can increase costs since cloud billing is designed for variable workloads, not steady-state origin workloads. 
  • Moving origin workloads to bare metal infrastructure enables streaming platforms to avoid high per-GB costs. 

Streaming platform engineers are often faced with the difficult challenge of reducing costs. Many concentrate on negotiating better content delivery network (CDN) contracts and optimizing the last mile of content delivery, for example by fine-tuning cache rules.  
 
But they are focused on the wrong part of the streaming process. Streaming costs are highest at where traffic leaves the origin. It’s at this origin layer that streaming platforms incur hefty egress fees from cloud providers. Delivering one petabyte of egress per month from a major hyperscaler can cost up to $50,000 in egress fees alone. That figure does not include compute, transcoding, or storage—it’s just the exit fee for data.  

The origin along with the cache shield generate the largest egress loads and therefore create the largest egress costs. By the time data reaches the CDN, those costs have already been incurred.  
 
To complicate the cost challenge, the origin and shield layers have sustained, predictable loads. They have constant throughput and constant egress—there’s no elasticity in their workloads. But that’s the worst possible profile for hyperscale cloud compute, which uses per-GB pricing to accommodate periodic bursts.  
 
So, to truly address streaming costs, streaming platforms first need to refocus on the layers that incur the greatest costs. They then need to rethink architectural choices that drive up those costs.  

Cloud Pricing Punishes the Origin Workload  

Cloud pricing is appropriate for workloads that burst. Elastic compute, machine learning training, GPU-based transcoding at scale, and similar workloads all experience temporary spikes. Paying per GB makes perfect sense: You only pay for what you use. 

By contrast, hyperscale cloud pricing is most expensive for sustained, high-volume egress workloads that do not burst—and these are the workloads at the origin. So, for example, software transcoding, live ingest, manifest packaging, and origin caching for adaptive bitrate (ABR) ladders are workloads that run continuously, without bursting. They never release resources from cloud environments. They push bytes at a constant rate as long as there are streaming platform users watching videos.  

For many streaming platforms, choosing bare metal infrastructure instead of cloud services for these workloads makes better economic sense. Organizations can incur a flat monthly rate for a committed bandwidth tier. Compared with the per-GB pricing model of hyperscalers, organizations can save tens of thousands of dollars even before any compute optimization is applied. When moving steady-state workloads from hyperscalers to dedicated hardware, organizations can reduce their total cost of ownership up to 40%.1 

Dedicated Bare Metal Controls Cost and Improves Output 

Organizations considering a move to bare metal for sustained workloads should look for vendors that clearly define vendor and streaming platform responsibilities. For example, Hivelocity provides single-tenant bare-metal hardware for ingest, transcoding, packaging, origin, and shield. But streaming platforms are responsible for codec choices, digital rights management (DRM) operations, CDN selection, manifest configuration, cache eviction policy, and the customer-facing delivery application programming interface (API). 

Establishing that clear boundary helps with the cost analysis, clarifying that the origin layer controls the egress economics. By contrast, the CDN can’t change the per-GB cost at which traffic left the origin.  

Using single-tenant bare metal changes the architecture in three important ways that benefit streaming platforms and their customers:  

  • Predictable costs: Flat monthly bandwidth pricing replaces the per-GB metering of cloud services. Organizations pay the predictable, committed costs for the origin and shield layers. They no longer pay variable, consumption-based exit fees. 
  • Increased quality: When organizations move steady-state workloads from the cloud to bare metal, they remove the need for CPU scheduling that can ultimately cause quality issues for transcoding and ingest packaging. CPU scheduling variance can generate encoding artifacts or dropped frames. Dedicated hardware removes those issues. 
  • Improved viewer experience: Dedicated hardware isolates the network path, eliminating the noisy-neighbor interference that can result in poor viewer experiences. That interference is particularly troublesome during live events, when egress volume spikes. 

Where to Audit Your Stack First 

If your organization is running your over-the-top (OTT) platform’s origin and shield workloads on hyperscale cloud, audit your stack to determine whether moving to a dedicated bare metal architecture will be beneficial. Start the audit with a few key questions. 

  1. What is your per-GB egress rate out of your origin region? And what percentage of total bandwidth costs does it represent? 
  2. Is your origin workload actually elastic? In other words, does compute usage vary significantly based on viewer count? Or (more likely) does it run at a sustained baseline regardless of concurrent audience size?
  3. Are your transcoding and packaging pipelines CPU-based software toolchains (such as FFmpeg, GStreamer, or equivalent vendor toolchains) running on shared virtual machines? If so, have you measured CPU scheduling variance against your encoding ladder consistency? 

The answers to these questions will tell you whether your egress costs are due to your CDN or your origin architecture. For most mid-market OTT platforms with petabyte-scale monthly delivery, the problem is with the origin. Rethinking that origin architecture can help organizations better align workloads with a pricing model that enables them to significantly reduce costs.  

Learn more about the benefits of dedicated servers for streaming platforms.

FAQ

Q: Why do egress costs start at the origin rather than the CDN?  
A: A CDN pulls content from the origin (either a cloud environment or dedicated hardware servers). The CDN copies content and stores it at the edge, close to users. But when the CDN needs a new file, or a file expires from the cache, it again pulls that file from the origin, incurring egress fees. 

Q: Why is bare metal infrastructure more cost effective for the origin layer? 
A: Hyperscalers charge egress fees by the GB. That’s fine for workloads that have occasional spikes. But it is more cost effective to use bare metal infrastructure for sustained high-volume workloads—like the ingest, transcoding, and packaging workloads that take place on the origin layer. Streaming platforms can just pay a flat monthly fee for egress. 

Q: Does switching origin infrastructure require changing CDN vendors?  
A: No. A shared-responsibility model enables streaming platforms to select their own preferred CDNs. The bare metal infrastructure vendor operates the ingest, transcoding, packaging, origin, and shield layer. The platform is responsible for the CDN routing logic, traffic shaping, and last-mile delivery. 

Citations:
1Datacenters.com, The Resurgence of Bare Metal Servers in 2025: Powering Performance, AI, and Cloud Alternatives, June 2025 

Come see what the Hivelocity difference
 means for your organization
Get expert guidance on choosing the right cloud solution for your enterprise needs.
Disaster Recovery
How to Survive When Ransomware Strikes
Don’t Miss What’s Next!
Register for live webinars, join expert AMAs, explore in-person meetups, and more.