
Excerpt: Two data-center proposals can quote the same megawatt capacity while describing fundamentally different amounts of usable AI compute. The difference is often hidden in power boundaries, PUE assumptions and excluded infrastructure.
The GPU-per-Megawatt Illusion: Why More GPUs on Paper Can Mean Less Compute in Reality
Two AI data-center proposals are placed side by side.
Both claim the same number of megawatts. One lists more GPUs, offers a lower cost per accelerator and appears to deliver superior density.
Then the assumptions are opened.
One proposal uses average accelerator consumption. The other evaluates the complete rack and the operating conditions the facility must actually support.
One applies an attractive PUE figure. The other separates total facility power from usable ICT capacity.
One reserves power and space for networking, storage, management, cooling support, maintenance and expansion. The other mainly counts GPUs.
The higher number may not represent better engineering. It may simply represent a narrower calculation boundary.
This is the GPU-per-megawatt illusion: the belief that the largest theoretical accelerator count inside a nominal power envelope must be the best AI data-center design.
It is not.
The relevant metric is deployable compute per megawatt: the infrastructure the complete facility can power, cool, connect, feed with data, maintain and operate reliably at the same time.
How More GPUs Appear on Paper
“GPUs per megawatt” is commercially attractive because it is easy to compare. The problem is that neither side of the ratio has a universal definition.
A megawatt may mean utility input, total facility power, data-hall power or usable ICT power at the rack. A GPU count may mean accelerator modules, complete servers, installed racks or fully operational cluster capacity.
That creates several ways for a proposal to look denser without necessarily delivering more usable compute.
| Category | Headline calculation | Complete engineering model |
|---|---|---|
| Power boundary | Treats most facility power as available to compute | Separates utility, facility and usable ICT capacity |
| GPU power basis | Uses average consumption or isolated GPU TDP | Uses complete server or rack input and a defined workload envelope |
| PUE | Uses a target, annual average or best-case value | Uses a stated design point and checks other operating conditions |
| Supporting ICT | Minimizes or excludes network, storage and management | Includes the systems required to operate the cluster |
| Operating state | Assumes every component is available | Shows capacity during maintenance and the stated failure condition |
| Reserve | Allocates every available kilowatt | Preserves justified operational and expansion headroom |
The issue is not necessarily false arithmetic. It is inconsistent scope.
A proposal does not become more efficient because network switches, storage arrays, control nodes, cooling-support loads or maintenance capacity have been moved outside the calculation.
Average Consumption Is Not Design Capacity
Average GPU consumption is useful for forecasting energy purchases, operating expenditure and expected utilization.
It is not automatically the correct basis for sizing every electrical and thermal component.
Large transformer-training clusters can move between workload states in near unison. Uptime Institute reports that this can create frequent step changes in power demand, sometimes every second or two. The size of those changes depends on the hardware, cluster configuration, workload and power-management settings. This does not mean that every GPU runs continuously at one theoretical peak. Nor does it mean that every facility should be oversized using one arbitrary multiplier.
A credible design evaluates:
- Maximum sustained workload
- Repeated power excursions
- Synchronization and workload diversity
- Manufacturer limits
- UPS, breaker and busway overload characteristics
- Cooling-system response
- Power-capping or smoothing policies
- Maintenance and redundancy states
- The required reliability level
Average power answers an energy question.
Design capacity answers an operating question.
The facility must support the agreed workload without relying on unplanned throttling, repeated overload or the later removal of racks.
PUE Is Not Electrical Headroom
The Green Grid defines Power Usage Effectiveness as:
PUE = total data-center energy ÷ ICT-equipment energy
PUE describes facility overhead relative to the IT load. It is not a generic safety factor, and it does not replace electrical coordination, transient analysis, redundancy planning or a complete ICT load schedule.
Formally, PUE is an energy metric. Applying a design-point PUE to power can be useful during concept-stage capacity planning, but it remains an approximation. Detailed design still requires component-level electrical and mechanical load schedules.
A measured annual PUE is also not necessarily the same as a target PUE, a design-point PUE or performance during the most demanding ambient and load condition.
Consider a deliberately simplified example.
A 10 MW facility allocation converted using a PUE of 1.10 appears to provide 9.09 MW for ICT equipment.
At a design-point PUE of 1.20, the same allocation provides 8.33 MW.
That is roughly 760 kW of difference before any allowance is made for networking, storage, management systems or operational reserve.
The example does not suggest that either PUE is correct for a particular project. It demonstrates why the basis must be disclosed.
ElioVP’s approach is not to claim that one “worst-case PUE” solves every engineering question. It is to use conservative and transparent design assumptions, then separately validate power distribution, cooling, workload behaviour, maintenance conditions and operational headroom.
GPUs Are Not the Complete AI Platform
A GPU does not train or serve a model by itself.
Useful AI infrastructure also requires:
- Host CPUs and system memory
- Scale-up interconnects
- Scale-out Ethernet or InfiniBand
- Fabric switches and optical transceivers
- High-performance storage
- Dataset-ingestion and checkpointing capacity
- Management and control-plane nodes
- DPUs and security services
- Monitoring and orchestration
- Customer and out-of-band networks
Liquid-cooled systems also require a complete thermal chain: rack interfaces, CDUs, pumps, facility-water systems, controls and external heat rejection.
Every one of those systems consumes power, occupies space, generates heat and adds cost.
A proposal that assigns almost the complete ICT envelope to accelerator power has not made the supporting platform disappear. It has left it outside the headline number.
A powered accelerator can still be commercially unproductive if the fabric is congested, storage cannot feed the workload, cooling limitations force power caps or the cluster cannot remain available during maintenance.
A powered GPU is not necessarily a productive GPU.
AMD Helios Makes GPU-Only Calculations Obsolete
AMD Helios is the clearest current example of why accelerator-only capacity planning no longer works.
The Helios rack-scale design integrates 72 AMD Instinct MI455X GPUs with AMD EPYC “Venice” CPUs, Pensando networking, UALink scale-up connectivity, rack-level power, liquid cooling and ROCm software. It uses the double-wide OCP Open Rack Wide format.
This is not a conventional rack containing 72 independent GPUs.
It is a coordinated rack-scale computer.
A credible Helios deployment must include:
- Complete rack input
- Host CPUs and memory
- Scale-up and scale-out fabrics
- External network-switch capacity
- DPUs and management infrastructure
- Storage and checkpointing
- Liquid-cooling distribution
- Residual room cooling
- Facility heat rejection
- Physical service clearances
- Maintenance isolation
- Software and power-management policies
In July 2026, AMD and Schneider Electric released a jointly developed Helios infrastructure reference design supporting rack densities up to 246 kW and modular clusters up to 10.4 MW of IT load.
The 246 kW value is a supported density in that reference design. It should not be interpreted as a statement that every Helios rack will continuously consume exactly 246 kW.
Its importance is architectural.
At that density, the electrical system, liquid-cooling plant, building geometry, controls and operational model must be engineered together.
A bidder cannot calculate credible Helios capacity by multiplying 72 accelerators by an individual GPU power figure. That would omit the host systems, fabric, networking, cooling and supporting infrastructure that turn those accelerators into a working AI platform.
AMD has announced that Helios shipments to customers, including Microsoft, will begin during the second half of 2026. The infrastructure requirement is therefore immediate, not theoretical.
NVIDIA Vera Rubin confirms the same direction.
NVIDIA describes Vera Rubin as a five-rack platform operating as one AI supercomputer. The platform combines NVL72 compute with dedicated CPU, storage, networking and operational infrastructure.
Different ecosystems are reaching the same conclusion:
The data center is becoming part of the computer.
Why the Realistic Proposal May Look More Expensive
A complete proposal may show fewer GPUs and a higher initial cost because it includes more of the real project.
It may include:
- Network fabrics
- Storage systems
- Management infrastructure
- Complete liquid cooling
- Heat-rejection equipment
- Maintainable power paths
- Monitoring and controls
- Commissioning
- Service space
- Expansion capacity
A narrower proposal can appear cheaper because those requirements have been excluded, deferred, assigned to the customer or left for detailed design.
The costs do not disappear.
They return later as additional switchgear, larger busways, extra network rows, storage upgrades, cooling-plant expansion, permanent power caps, rack depopulation or delayed commissioning.
The cheapest proposal is sometimes simply the proposal with the most costs still hidden.
Headroom is not automatically waste.
It may be required for workload changes, maintenance, demanding environmental conditions, partial cooling availability, additional networking and storage, commissioning uncertainty or future rack generations.
Excessive reserve can also strand capital. The answer is not blind oversizing.
It is evidence-based engineering with clearly disclosed assumptions.
What Buyers Should Ask Before Comparing Proposals
Before accepting a GPU-per-megawatt figure, buyers should ask:
- What exactly does the quoted megawatt represent?
- Is GPU power based on component TDP, average consumption, complete server input or complete rack input?
- Is the PUE annual, measured, targeted, seasonal or a defined design point?
- Are networking, storage, management and cooling-support systems included in power, space and cost?
- What capacity remains during maintenance or the stated failure condition?
- How are sustained workloads, recurring power changes and power-management policies evaluated?
- Can the quoted configuration actually be commissioned and operated without unplanned derating?
If bidders cannot answer those questions using equivalent boundaries, their GPU-density figures should not be compared.
ElioVP Designs for Deployable Compute
At ElioVP, capacity planning begins with the complete AI platform rather than an isolated accelerator specification.
That means connecting facility power, liquid cooling, rack architecture, networking, storage, management infrastructure and workload behaviour in one capacity model.
It also means making assumptions and exclusions visible before contract signature, not after commissioning.
ElioVP’s capabilities span ModFlex modular data centers, high-density hardware, storage, advanced networking and AMD workload optimization through Paiton.
That combination is particularly relevant to Helios.
Being ready for AMD Helios is not only about providing enough electrical power or installing liquid-cooling pipes. It requires an understanding of the relationship between MI455X compute, EPYC host systems, UALink, Pensando networking, storage, ROCm, workload behaviour and the physical facility.
ElioVP’s engineering team also includes an Uptime Institute Accredited Tier Specialist, bringing formal Tier Standards knowledge into decisions concerning resilience, maintainability and operational requirements.
Being ready for AMD Helios and NVIDIA Vera Rubin does not mean claiming that every final platform can be installed in any building without project-specific validation.
It means being prepared to engineer the complete rack- and pod-scale system around the real site, workload, cooling architecture, power topology, regulatory environment and expansion plan.
At ElioVP, we would rather explain why a realistic number is lower today than explain why an unrealistic number cannot be delivered tomorrow.
Productive Compute Is the Metric That Matters
AI data center power requirements cannot be reduced to the largest GPU number that fits into a spreadsheet.
Average power is not design capacity.
PUE is not a safety margin.
Accelerator TDP is not complete rack input.
Installed GPUs are not necessarily deployable GPUs.
Do not ask only how many GPUs fit inside a megawatt.
Ask how many GPUs the complete facility can power, cool, connect, feed with data, maintain and operate reliably, at the same time.
That number may be lower than the most aggressive headline.
It is also far more likely to survive detailed engineering, commissioning and real workloads.
ElioVP can independently review an AI data-center capacity model, normalize its assumptions and identify which requirements have, or have not, been included before a theoretical GPU count becomes a costly physical constraint.
Frequently Asked Questions
How Many GPUs Can an AI Data Center Support per Megawatt?
There is no universal number.
The result depends on the accelerator architecture, complete server or rack configuration, facility-power boundary, PUE, workload, networking, storage, cooling, maintenance conditions, redundancy and justified reserve.
For rack-scale platforms such as AMD Helios, capacity should be calculated using the complete system and its supporting infrastructure, not by dividing one megawatt by the power rating of an individual GPU.
Should an AI Data Center Be Sized Using Average or Peak GPU Power?
Neither value is sufficient by itself.
Average power is useful for estimating energy use and operating cost. Engineering design must also evaluate maximum sustained demand, recurring excursions, workload synchronization, equipment limits, power-management policies and the required reliability objective.
The correct basis is a validated operating envelope rather than one average or one theoretical maximum.
Does PUE Include Networking and Storage?
Networking and storage belong inside the ICT load. They are not optional facility overhead that can be ignored when calculating available GPU capacity.
PUE describes the relationship between total facility energy and ICT-equipment energy. It does not determine how the ICT envelope should be divided between GPUs, CPUs, networking, storage and management infrastructure.
Why Do GPU-Capacity Estimates Differ Between Proposals?
They often use different definitions.
One bidder may start with utility power while another begins with usable ICT capacity. They may also use different PUE assumptions, GPU power bases, workload profiles, support-system allowances, maintenance states and expansion reserves.
The proposal with the higher GPU count may be more efficient, or it may simply have counted less of the complete platform.
What Does “Helios-Ready” Mean for a Data Center?
A Helios-ready design must consider the complete rack-scale architecture: GPUs, CPUs, UALink connectivity, Pensando networking, rack power, liquid cooling, external fabric, storage, management and ROCm operation.
It must also accommodate the double-wide Open Rack Wide format, service access, facility-water interfaces, heat rejection and the required maintenance conditions.
It should not mean that a generic rack position has simply been labelled “AI-ready.”
