AI data center power architecture: A/B Feeds, UPS and Rack PDUs
E
ETENZ•Editorial Team
AI data center power architecture must hold up when one feed is lost and equipment is taken out for maintenance. Start at the server power inputs, then size the A/B paths, UPS backup, rack PDUs and metering around the required operating state.
Two servers can both have dual power feeds yet behave very differently when one feed disappears. One may continue at full load, another may throttle, and a third may shut down. That behaviour determines the rating of each distribution path, the UPS arrangement and how much compute capacity remains available during a fault.
Power design usually starts with a single-line diagram. For an AI compute module, that diagram should extend to the server power inlets and match the required response after a feed failure: which workloads keep running, which may be reduced, and which can stop after saving state. Set those rules before the RFQ, and the trade-offs between a single feed, A/B feeds and different UPS arrangements become much easier to compare.
Start with the server power inputs
Take the NVIDIA DGX H100/H200 as an example. Its six power supply units are arranged for 4+2 redundancy, and at least four must remain powered for normal operation. If three PSUs are connected to feed A and three to feed B, losing either feed leaves only three. The system may then run at reduced performance or shut down. NVIDIA's data center design guide for DGX SuperPOD therefore treats conventional A/B power as suitable for management racks only. Where availability is critical and a workload cannot recover from a checkpoint, it recommends at least three independent feeds per rack, with two of the server's six PSUs on each feed. Four PSUs then remain powered when any one feed is lost or isolated for maintenance.
The way PSUs are grouped determines both the number of feeds and the capacity required on each one. Put compute, storage and network equipment in the first load schedule. For every device, record the input voltage, plug type, number and redundancy of PSUs, intended feed grouping, and whether it continues normally, throttles or stops after one feed is lost.
A single-cord device that must draw from two sources needs a rack transfer switch, with a transfer time shorter than the equipment's ride-through time. Dual-cord devices should connect to separate A and B rack PDUs in line with the manufacturer's instructions. Plugging both cords into the same PDU may look redundant, but it still leaves a single power path.
The business requirement decides which loads must stay live. Can training jobs recover from checkpoints? Can inference move to another node? How long must storage and the management network remain available? Once the business and IT teams answer those questions, the electrical design can define the backed-up load.
Size each A/B path for the required fault state
A single distribution path has to be shut down for maintenance, and a fault anywhere along it interrupts the load. It suits loads that can stop or workloads that another node can take over. With A/B feeds, trace both paths upstream from the rack to their sources, distribution busbars and protective devices. If both feeds ultimately depend on the same UPS or bus section, that shared point can take them down together and should be shown clearly on the single-line diagram.
Capacity is set by the condition after one path is lost. Suppose an IT load requires 100 kW and its power supplies share the load normally but allow either path to carry the full load after a failure. In normal operation, A and B each carry about 50 kW. If A fails, B must carry 100 kW. Two paths rated at 70 kW each look comfortable in normal operation but are insufficient with only B remaining.
Run that calculation through every element of the surviving path: UPS output, distribution circuit, busway or cable, rack PDU and outlet, all the way to the server inlet. The lowest-rated element sets the usable capacity of the entire path. Allowance is also needed for ambient derating, future expansion and load variation. In the example above, two paths with a combined nameplate capacity of 140 kW can support only 70 kW of IT load if either path must carry the full agreed load.
Path A out: B carries all 100 kW
If the project accepts load reduction after a feed failure, state the reduced load, the time allowed to reach it and the control method in the load schedule. During the interval between the failure and completed load shedding, the surviving path and its protective devices still see the full load.
Equipment that can use three feeds, with two PSUs on each as in the DGX example, needs less capacity per path. For the same 100 kW load, the three feeds carry about 33 kW each in normal operation. After one is lost, the other two carry 50 kW each, so each path can be sized for 50 kW and the three total 150 kW. A mutually redundant two-feed arrangement needs 100 kW on each path, or 200 kW in total. The three-feed option reduces capacity per path but adds a third distribution route and another set of connections.
Choose which path the UPS protects and where it sits
A UPS on a single feed can bridge an upstream outage for as long as its battery supports the load, but there is still only one downstream distribution path. In an arrangement with a UPS on feed A and utility power directly on feed B, a utility outage transfers the full load to A. The rating of path A, the server's available power with one feed, and the battery runtime must all be checked for that moment.
A separate UPS on each of the A and B paths gives both feeds stored-energy backup. Maintenance still depends on the actual bypass, isolator and battery connections: where the bypass supply comes from, what loses protection during maintenance, and how much load the remaining path can carry. When a UPS is placed on maintenance bypass, the load remains energised from utility power, but that UPS battery no longer protects it.
In many quotations, N+1 means one extra power module inside a modular UPS. One module can fail without reducing the output, but that says nothing about redundancy in the output bus, bypass or downstream circuits. Two systems both described as N+1 can therefore provide very different remaining capacity during maintenance or a fault. The quotation should identify exactly which component is redundant.
UPS and batteries may sit inside the compute module, in a separate power E-House, or in an existing site UPS plant while the module contains distribution only. Installing them in the compute module uses rack space and adds their losses to the internal heat load. A separate power E-House leaves more room for IT racks, with the A and B paths brought across by separate cables or busways and the interfaces matched on both sets of drawings. Battery location also affects ventilation, service access and fire compartmentation.
Servers continue producing heat while running on UPS power. In a liquid-cooled compute module, the backed-up load schedule should therefore include CDU circulation pumps, controllers, critical network equipment and any essential air-cooling equipment alongside the IT load. Whether the outdoor heat-rejection equipment also needs backup depends on the permitted temperature rise, available thermal storage and the time needed for an orderly shutdown. Backing up only the controller can leave it reporting alarms after the circulation pump has already stopped.
Battery runtime is meaningful only with its load conditions. A stated 10-minute backup can produce very different results depending on the kW load, the equipment covered, battery temperature, whether capacity is assessed when new or at end of life, and the end voltage. The supplier should provide discharge data for the specified configuration and conditions. If the battery bridges the start of a generator, include generator start-up, transfer and load-acceptance time.
Select rack PDUs for fault-state current
Outlet count is only one part of rack PDU selection. The input phase arrangement and current rating, protection for each outlet bank, and the ratings of outlets and power cords must all match the server connection plan. List rack PDUs separately from upstream distribution boards in the RFQ.
A simplified example shows why the fault state matters. For a rack with 60 kW AC input at balanced three-phase 400 V and a power factor of 0.95, the line current is I = P / (√3 × U × power factor), or about 91 A. With an even A/B split, each path normally carries about 46 A. If either path must support the full 60 kW after the other is lost, the surviving path carries about 91 A. A 63 A input PDU on each path is adequate in normal operation but exceeds its rating in that fault condition.
Final selection must also consider the most heavily loaded phase, each outlet-bank circuit, cable installation conditions and equipment derating. Total kW may show spare capacity while one phase or outlet bank reaches its limit first. The permitted use of a continuous-load rating varies with the applicable project requirements and component conditions.
GPU demand moves with the training workload and can change sharply over a short interval. Supply the size and duration of those steps, together with the power-up profile, so the UPS dynamic response and protection coordination can be checked. Monthly energy data and one-minute average power do not reveal these short transients.
Place meters along the power path
A meter at the module incomer shows total module consumption. Meters at the UPS input and output reveal conversion losses and the load on the UPS output. Rack PDU metering supports capacity allocation, phase-balance checks and decisions about adding equipment. Put cooling on separate circuits so IT and auxiliary consumption remain visible as separate figures.
For dual-cord IT equipment, add the A and B readings over the same time interval and at the same measurement boundary to obtain its energy use. An upstream meter already includes downstream consumption, so adding both levels counts the same energy twice. If the readings are used for billing, agree the meter accuracy, calibration and settlement method separately.
A monitoring point list will usually include current, power, energy and power factor by phase, switch status, UPS operating mode—utility, battery or bypass—and remaining battery capacity. When one feed fails, operators need to see the event and the corresponding rise on the surviving path in the same view before deciding whether to shed load. Set the sampling interval, event timestamps and recovery of data after a communication outage around alarm and post-event analysis needs.
Test transfer, maintenance and recovery
Build the acceptance plan around operating scenarios: normal load, loss of A, loss of B, UPS battery operation, maintenance bypass and restoration of power. For each test, state the starting load, the point to be opened, the expected response, the permitted interruption or load reduction, and the voltage, current, alarms and workload state to be recorded.
The phrase ‘open feed A’ is not precise enough. Opening the UPS input transfers the UPS to battery and the server still receives that feed; this tests the battery. Opening the UPS output or the rack branch removes one server feed and tests whether the other path can carry the entire load.
Distribution paths, capacity, controls and metering can be tested at the factory with simulated loads. Server throttling, network reconnection and workload interruption can only be observed on site with the actual equipment connected. Protection coordination is confirmed through design checks and the relevant tests.
Turn the design into a prefabricated module
Inside the module, these requirements determine cabinet layout, busway and cable routes, battery location, heat removal and maintenance clearances. ETENZ can manufacture the E-House enclosure and, according to the project split of responsibilities, integrate electrical, thermal-management and monitoring systems. Delivery options include prefabricated interfaces, factory assembly and complete module delivery, with OEM/ODM cooperation available.
For an RFQ, extend the single-line diagram to the server power inlets. Show how each device's PSUs are grouped and how it responds to the loss of one feed. Include cooling equipment in the backed-up load schedule, define the acceptance scenarios, and attach the available site power conditions. Suppliers can then show power paths, capacity and backup scope on the same drawing, making competing quotations much easier to compare. The DGX H100/H200 power details in this article come from the power specifications in the NVIDIA DGX H100/H200 User Guide and the electrical specifications chapter of DGX SuperPOD: Data Center Design Featuring NVIDIA DGX H100 Systems.
Tags
AI data center power architectureA/B power feedsdata center UPS designrack PDU sizingmodular data center power