Power11 has finally been launched, and what we’re seeing is an evolution, not a revolution, in silicon. It’s a very similar chip to Power10, but an updated 7nm process and all-new packaging technology means we’ll get more threads (up to 240 per socket from 192), higher clock speeds, and higher memory bandwidth (P10 peaked at 800 GB/s, P11 is up to 1.2 TB/s). Since P11 is designed for large, multi-socket configurations, this makes it a very capable machine for all kinds of computations, including HPC. And, of course, Spyre is coming to the platform, which we’ve already talked about here.
Much of what IBM is doing with Power11 revolves around the administration software. I know, if you’re an HPC user, nothing could be more boring, but the reality is that downtime is everyone’s problem, not just IT’s. If you’re an HPC user, think of all the times your job got killed early or you couldn’t get anything done because the cluster was down for maintenance. Annoying, isn’t it? Wouldn’t it be nice if the cluster was online 24/365? Power11 answers those questions, because IBM’s pushing new limits with the P11 sysadmin stack to ensure that users experience zero planned downtime.
An important new step is providing all the software tools to enable large deployments to update system software without taking the machine offline. As somebody who has experienced an HPC system going offline for a full week for maintenance, zero planned downtime sounds impossible, but for very experienced Power admin, it’s something they already know how to do. But now, IBM has provided an automated toolkit so that anybody who has a P11 system can ensure they never have to take the system offline.
Now, how does IBM do it? The secret lies in PowerVM. If you’re used to x86, which let’s be frank, HPC in 2025 is almost entirely x86-based, you’re used to running bare metal due to the immense cost and performance overhead of traditional VM applications, so you may only be dimly aware of what VMs can do. Among other things, a VM allows you to move a running workload to another machine without stopping it. This is what allows system upgrades with no downtime.

Power is unique in that the PowerVM hypervisor is built into the firmware, so there’s no actual difference between “virtualized” and “bare metal.” Your operating system is always virtualized, even when you have a single logical partition you never touch, without the traditional overhead third-party solutions impose on other systems. This means if your cluster is built out of Power, you can deploy firmware updates and other upgrades simply by moving workloads around while you do maintenance, rather than taking the server offline. This has always been possible with PowerVM, but you had to know what you were doing. As of Power11, the process is now fully automated.
Surprised? You shouldn’t be. IBM invented the hypervisor way back in the 1960s. If you think HPC systems are intolerant of unnecessary overhead, try keeping software lean enough for a fussy old System/360 with only 512 kB of system memory. 60 years of thought leadership is why PowerVM is the only full-stack integrated hypervisor in the server industry.
Is there more? Of course there’s more. IBM Concert is a powerful new tool that uses AI to identify operational risks, provide actionable insights, and automate remediation, including deploying patches. Combined with zero-downtime technology, this means your HPC system can stay continuously updated and secure, rather than only being maintained 1x-2x year, as is common in the industry. And as we’ve talked about before, a secure HPC system is one that can be fully integrated into your enterprise infrastructure, eliminating the frictional losses typically associated with the extreme measures taken to quarantine typically insecure HPC systems from the rest of the network
And it’s all due to the fact that IBM Power has virtualization embedded at the lowest possible level.