It wasn’t that long ago that industry analysts were predicting the death of the on-premise server and the migration of virtually all enterprise computing to “the cloud.” Somewhat famously, Gartner Analyst Dave Cappuccio predicted in 2019 that by 2025, 80% of all businesses will have shut down their data centers and moved everything to “the cloud.” I put “the cloud” in quotes because there is, in fact, no single cloud. There are lots of clouds offered by many companies, and there’s nothing unified about them. You’re just paying somebody else to host your workload on their computer.

Of course, it’s 2025, and if people are supposed to be shutting down their on-premise servers, nobody told our customers, or countless other business whose day-to-day computing needs range from a single-core inventory server to an energy-hungry GPU cluster for HPC-AI workloads. Cloud ends up being right for many situations, but not all, and that’s why we’re starting to see a significant amount of repatriation.
While there are a lot of technical reasons on-premise and edge computing make more sense than cloud, one major reason is simple economics of buying vs renting. Now, the way accountants run their figures, renting may seem like a slam dunk over buying. Let’s say we have a $1,000,000 cluster (CPU or GPU, doesn’t matter for this exercise) on a 5-year depreciation schedule, and suppose we paid cash up front. This will show up on our books as a $200,000/yr expense. Additional costs from hosting and IT might double or even triple that. Let’s say it ultimately comes to $500,000/yr for the whole package. The way the spreadsheets work, if I can negotiate a contract for $400k/yr from a cloud service for the same amount of compute, switching is a no-brainer. But there are several hidden costs here:
#1 You can’t actually fire your IT.
HPC-AI IT workers are valuable professionals who are deeply embedded in your business processes. They’re not just there to make sure the lights stay blinking on the servers. Most of their work is spent managing the software stacks and user needs associated with the system, things unique to your business that a cloud service won’t provide to you for free. Oh, sure, they might offer IT as a service too, but that price tag isn’t going to be small, because HPC-AI IT isn’t a 2 hour a week job.
#2 HPC-AI systems rack up a lot of hours
If you are paying by the hour, your typical cloud HPC-AI price with a multi-year contract is around 3x as expensive as your on-premise system. This means that if your HPC-AI machine is getting used over 30% of the hours in the year, and many of these machines run hot for 75%-90% of the year, you’re not going to save money by going cloud-native.
#3 You can’t save by delaying a replacement
And lastly, we come to the big one, something that affects everything from delivery trucks to engineering clusters. When you spend money on capital equipment, if you run into a money crunch, you can cut a lot of costs by keeping that equipment past its depreciation date. If you keep that machine for a 6th or even 7th year (and trust me, in 6 years, today’s $1m machine will still have a hefty amount of compute compared to a workstation), it only keeps costing you IT & hosting. That $200K item on your books from depreciation disappears. By contrast, you simply can’t zero out any part of your cloud cost. You can only buy fewer hours next time around.
Hybrid: The best of both worlds
And yet, there are still fundamental issues with getting all the compute you need on-prem. These days, even a consultant working alone can easily generate a huge mesh that can require $500K or more worth of hardware to simulate quickly. This is why we’re seeing more and more HPC users turn to hybrid computing. Perhaps you have a 1024-core cluster on-premise for your day-to-day computing, but occasionally, for those really big jobs, or when you just have more jobs than your machine can digest, you burst out to a cloud service. This kind of arrangement lets you easily manage the costs of cloud computing, benefit from the economics of on-premise, and still get access to the really big machines for that 5% of the year that you actually need them. And if you do run into a cash crunch, you can zero out cloud completely and still have HPC capability on hand.