Industry Insights

Six interfaces to settle before a liquid-cooled AI module RFQ

E
ETENZEditorial Team

Quotations become comparable once the liquid cooling interfaces are written down. Start from the server platform, then fix CDU scope, duty points, flow balance and pressure ratings, coolant chemistry, and what happens when cooling is lost.

Conceptual cutaway illustration of a liquid-cooled compute module, showing server racks, a CDU, internal pipework and outdoor heat rejection equipment

Three quotations for the same liquid-cooled compute module can cover three different sets of equipment. One stops at a supply and return connection on the enclosure wall. One adds the CDU and the rack manifolds. One carries the outdoor heat rejection as well. Those three prices are not comparable, and the reason is scope rather than price. Interfaces become comparable only when each one is fixed to a drawing, to a duty point and to a named responsible party, which is also what turns the site connection into something a programme can be built around.

This article starts from AI servers on single-phase cold plates and works through the six interfaces worth settling before an RFQ goes out. Part of the engineering logic below is drawn from Chinese national and association standards for cold-plate liquid cooling, which are among the few published documents that put figures on this equipment; the sources are named at the end. Two points about that. GB/T 48023-2026 is a recommended Chinese national standard and takes effect on 1 February 2027, so it is still in its transition period. And "shall" inside a document of that kind is that document's own requirement, not a legal obligation, in China or anywhere else. For a project in any other country the binding rules are the ones in force at the project location. That is exactly why an RFQ has to name the standards, editions, scope and acceptance criteria the project will be held to, instead of leaving "compliant with the applicable standards" to be interpreted at delivery.

1. Servers and racks: split the heat load before anything else

Set out the heat load for one stated configuration: total server power, how much of it the cold plates remove, and how much is left to the air inside the module. Putting the GPUs on liquid does not stop power supplies, networking and drives rejecting heat to the room, and the size of that remainder decides whether the module needs in-row units, a fan wall or nothing at all. One part of the cooling load never appears on a server datasheet: the pumps, the pipework and the buffer volume add heat to the coolant themselves. Ask whether the design cooling load you are quoted already includes that circulation gain, and state the configuration every figure refers to, so that two bidders are not answering different questions.

A first round of enquiry does not need the final platform. Candidate server models, rack quantities and dimensions, the intended build configuration and the phased expansion plan are enough to start. Rack dimensions are the one item not to fill in from air-cooled habit: liquid racks tend to be wider and deeper, and laying them out at air-cooled pitch produces rack positions that will not exist. Ask each server vendor for the interface document, covering permitted coolant inlet temperature, required flow, pressure drop, maximum allowable working pressure, fluid and connector type, because those figures are the inputs to everything in the next five sections. Where the platform is undecided, comparing two or three candidates side by side gives a more usable layout than one averaged rack-power figure applied to the whole room.

Geometry is where borrowed numbers do the most damage. Clear height under a raised floor, aisle clear width and floor loading are quoted freely in this industry, but the published figures come from documents written around one country's practice and most of them carry conditions; the commonly cited floor void, for instance, assumes the liquid lines share that void with power and data trunking, and stops applying the moment the electrical services go overhead. Rather than import a number, state the routing decision and derive the space from it. Where do the headers run, under the floor, overhead or along the sides? Are liquid and electrical services kept to separate aisles? May a header pass directly above or below a rack? What clear width is needed to withdraw a server and operate a quick disconnect? What is the floor loading with the system filled? Those are the items the enclosure designer needs on a drawing, and they are the items a quotation can be held to.

2. CDU and pipework: put every handover on a numbered fitting

A liquid-to-liquid CDU keeps two circuits apart with a heat exchanger. The facility side (FWS) carries heat out to the outdoor heat rejection or the site cooling plant; the IT side (TCS) serves the rack manifolds and the cold plates. Each side has its own circulation, and the two fluids never meet. The arrangement is simple; the naming is not. "Primary" and "secondary" are used in opposite senses by different manufacturers and different guidance documents, and one word written the wrong way round swaps an entire set of conditions. Settle it on the drawing: label each circuit by name, mark the flow direction, and identify supply and return at every connection. The Open Compute Project's Data Center Liquid Distribution Guidance and Reference Designs is worth reading alongside this for circuits, interfaces, maintenance and loss-of-cooling scenarios. It is design guidance, not a requirement, and it does not replace the specification the project still has to write.

Put the supply list on one page and give every line an owner: outdoor heat rejection, facility-side pumps, CDU, module headers, rack manifolds, hoses and quick disconnects, coolant, instrumentation. Who supplies it, who installs it, who fills it, who commissions it. Two maintenance decisions on that list reach back and change the layout inside the module, so they belong in the RFQ rather than in a site query six months later. Decide whether the filters and the vent and drain devices have to be serviceable while the system keeps running, and say so explicitly. Running means the system stays in operation, not that anything is opened under pressure: the section being worked on is isolated and depressurised first, which only works if a standby path can carry the required flow while that section is out. A duplex filter arrangement is the usual answer. Isolation valves either side of a single filter are not, because closing them takes that supply path out of service. Whichever concept is chosen, the module has to reserve space for the valve groups and room for someone to stand and operate them.

Every handover should land on a numbered flange, quick disconnect or terminal, with the mating half's supplier, the connection specification, the seal material and the position all written down. Which jointing methods are acceptable, whether welded, flanged, grooved or threaded, depends on line size and on the code applicable at the project location, and threaded joints are commonly restricted to the smaller lines. State the permitted methods by size, name the code the fabrication will be held to, and say which joints must be made in the factory. Prefabrication makes that worth doing: large headers welded and tested on the production line are a different proposition from the same joints made on a live site. Weld procedure qualification and inspection requirements follow the same rule, set by the applicable local code and the project specification rather than assumed. External port coordinates, transport caps and the site connection sequence then only have to agree with the general arrangement drawing.

Concept diagram of two separate liquid cooling circuits. On the facility side (FWS), a circulation pump drives fluid from the dry cooler through the facility channel of the plate heat exchanger and back to the dry cooler. On the IT side (TCS), the CDU pump drives coolant out of the IT channel of the exchanger into the blue supply manifold, through the server cold plates, and back along the orange return manifold to the exchanger. The two fluid channels are isolated and transfer heat only.
Concept: blue is supply, orange return. FWS and TCS trade heat across the exchanger; fluids stay apart.

3. Cooling conditions: compare capacity at one duty point

A CDU rating and a heat-rejection rating mean little until both sit on the same duty table. The RFQ should state the server's permitted coolant inlet temperature, the design supply and return temperatures on both sides of the CDU, the approach temperature across the heat exchanger, the flow rate, fluid and concentration on each side, and the design ambient conditions, and should then require performance data at the target load, fluid and flow rather than a nameplate figure. Allowable coolant temperatures are a product characteristic: where a standard publishes ranges they are recommendations, and specific products are qualified to their manufacturer's own limits. Condensation control is set separately, against the dew point of the environment the equipment actually sits in, not against a supply temperature carried over from another project.

With a dry cooler the temperature chain runs outward from the servers. The permitted inlet temperature at the cold plate sets how much approach the CDU may consume, which sets the facility supply temperature required, which decides whether a dry cooler can deliver it at the site's summer design dry-bulb temperature. How closely a dry cooler approaches ambient is a selection result that moves with coil size, air flow, fluid and load, so treat any rule-of-thumb margin as a direction check and put nothing into a quotation that does not come from selection data at the target duty. The direction is still worth knowing: the hotter the design ambient, the warmer the water available, and in hot conditions a dry cooler on its own frequently cannot reach the inlet temperature the servers allow. Whether supplementary cooling is needed, by adiabatic or spray pre-cooling, a cooling tower, or mechanical cooling in parallel, follows from whether the heat-rejection capacity meets the requirement at the design condition. Capacity is the test, not whether the site is labelled a hot-climate location.

One selection input is missed often enough to be worth naming: the weather basis itself. State which station's data is used, over what record period, whether extreme dry-bulb and wet-bulb values or a design percentile are applied, and whether any uplift is added for heat-island effects and for recirculation between units on the site. The outdoor equipment is sized on that number, and a bid using a milder one will always look better than it is. Inside the module, dew point and the surface temperature of cold pipework and fittings set the insulation, the condensation detection and the control response. The air-cooling equipment needs its own stated duty as well, or the heat that never reaches a cold plate ends up with no owner. Ask every bidder for available capacity and operating limits at those stated conditions.

4. Flow and pressure: check balance and pressure rating separately

Where the CDU carries its own circulation pumps, the differential pressure it offers at its external connections has to cover, at design flow, the whole of the most unfavourable path: headers, valves, filters, manifolds, hoses, quick disconnects and the servers themselves. If a supplier quotes total pump differential pressure, the CDU's internal losses come off that before anything is allocated to external pipework, so ask for the external figure explicitly. Ask as well for the flow range over which the IT side can be controlled and the steady-state tolerance the equipment holds on supply temperature, then check the first phase against it. A hall built in stages, with half the racks installed, can sit below the range in which the equipment is able to control at all. That is a question to resolve before the order, not during commissioning.

"Total flow is sufficient" and "every server has enough" are different statements, and only the second one matters. GB/T 48023-2026, the Chinese national standard for cold-plate liquid cooling systems in data centres that takes effect on 1 February 2027, makes balance measurable with three limits: across the outlets of one distribution manifold, the spread between the highest and lowest flow should stay within 10% of the average of those outlets; between the branches of the IT side, the same 10%; and cold plates of the same specification should be within plus or minus 10% of each other on flow resistance. Like all GB/T documents these are recommended rather than mandatory, and they were written for cold-plate systems of this kind. For a project elsewhere they are a well-defined and comparable reference, which means that if you want them, they go into the RFQ as stated acceptance criteria rather than travelling as an assumption. One caveat comes with them: balanced is not the same as sufficient. Branches can be beautifully even and uniformly low together. So alongside the three balance figures, require written confirmation of the flow each rack and each server actually receives at the design condition. Both are things a supplier can put in writing before the order is placed.

Pressure ratings are checked component by component, not as a single system number. Ask for a schedule of every pressure-containing part, covering quick disconnects, manifolds, hoses, valves, the heat exchanger and the cold plates, with the maximum allowable working pressure each manufacturer states for it, because the allowable pressure of the whole circuit is set by the lowest entry on that list and it is often a connector rather than a pipe. Keep test pressure and maximum allowable working pressure as two separate numbers: a proof test is run at a multiple of the rating in order to verify it, and that multiple cannot be turned around to justify operating a part continuously at a higher pressure. For each test, agree the object under test, the medium, the pressure, the hold time and the pass criterion, and record them as their own line item. Add a node-pressure schedule for running, pump-off and thermal-expansion cases including static head, and fix the measurement points in the RFQ, keeping gauge, absolute and differential pressure distinct.

5. Coolant and pressurisation: cover first fill and maintenance too

The coolant and every wetted material, from cold plates and heat exchanger to pipework, seals and quick disconnects, form one system, and the variable that governs it is the cold-plate metal. T/CECS 1722-2024, a Chinese association standard and a recommended document rather than a mandatory one, splits the pH window it recommends by material: 8.0 to 10.0 for copper cold plates and 7.5 to 9.5 for aluminium ones, which leaves a common band of only 8.0 to 9.5. It separates the corrosion inhibitors the same way, recommending organic-acid salts for aluminium systems and azoles for copper alloys. On aluminium and copper sharing one coolant circuit, that document advises against it because of the corrosion risk on the aluminium side; it does not prohibit it. Those figures apply to systems within that standard's own scope, so on a project anywhere else they carry weight only if the RFQ adopts them explicitly, which is also the best use for them: a stated, checkable limit instead of an assumption. All of which makes the cold-plate metal an interface condition rather than a detail. Name it in the RFQ together with the coolant and its working concentration, list the complete wetted-material set, and ask for written confirmation of compatibility covering that list, along with the chemistry limits, the sampling interval and the replacement interval. "Deionised water" on its own is not a coolant specification.

Folklore about pipe materials travels further than the documents behind it, and outright bans get attributed to standards that never wrote them. What an RFQ needs is narrower and more useful: a requirement that every wetted material is confirmed compatible with the selected coolant, the pipe material and grade actually being offered and priced, and explicit rules for jointing and sealing workmanship. One rule is worth carrying across from the Chinese national standard already mentioned, which writes it as a requirement of the document rather than a suggestion: PTFE tape and liquid thread sealant are not to be used on coolant joints. That is the kind of rule a works procedure can apply and an inspector can check joint by joint at factory acceptance, which is more than most material clauses ever achieve.

Pressurisation needs three things fixed: the capacity of the expansion device, its pre-charge pressure and where it connects, all derived from the fluid volume, temperature swing and node pressures of the circuit it serves. Then say where make-up and venting happen and who owns them, and how overpressure protection is set. Where heat exchangers and isolation valves divide a circuit into sections, confirm each section on its own, and include any length of liquid that becomes trapped when a valve closes. The quotation scope should also name factory cleaning, transport preservation, site flushing, first fill, repeat sampling and consumables: state the flushing medium, the circulation time and the endpoint criterion, and require every filter element to be replaced before the working coolant goes in. Leak and pressure testing is agreed the same way, per test, by object, medium, pressure, hold time and pass criterion. The pressurised leak-test procedure often quoted for cold-plate heat exchangers and their manifolds comes from a Chinese association standard as a recommended method for those components specifically. It is not a general acceptance method for a whole circuit or a whole module, and the module-level test has to be agreed on its own terms.

6. Loss of cooling: turn alarms into defined actions

Loss of facility cooling, a CDU pump stopping, low flow on a branch, a leak, and loss of communication are five different events, and each one gets its own row: the detection point, who receives the alarm, what transfers or isolates, and at what point the IT control system is asked to shed load or shut down in a controlled sequence. Leak handling is where scope quietly goes missing. Detection, alarm and leak-location are usually specified, including for the facility-side pipework entering the IT space, while automatic shutoff frequently is not, and its absence from a standard's list means neither that a project can do without it nor that it is an optional extra. Decide whether a valve closes, who closes it and on what condition, then write that into the interlock table and into the scope of supply. Assuming the other party has included it is how a module arrives without one.

"Connects to the monitoring platform" only becomes a deliverable as a point list plus command permissions: which values are read-only, which commands may be written, who issues them, what state the equipment holds when communication is lost, and how it restarts when communication returns. Map that across the CDU, the module control panel and the project monitoring system, because those three are rarely from the same supplier. One monitored value bears directly on the enclosure design, and that is dew point. It calls for representative temperature and humidity sensing inside the module, and where those sensors sit decides whether the anti-condensation response fires when it should or when it should not, which makes it an item for the enclosure designer to settle on the drawing rather than for the site to discover. The same goes for drip trays under the pipework and under the CDU and whether they drain anywhere: containment and drainage sit in the enclosure scope, so the delivery split has to say so and the post-installation checks have to include them.

Before discussing how many minutes a system rides through a loss of cooling, put the parameters underneath that claim on the table: the failure scenario, the circulation that remains, the fluid volume actually taking part in heat transfer, the allowable temperature rise and the load-reduction ramp. Once circulation stops, the volume sitting in a remote tank cannot be counted as cooling still available to the servers, whatever its size. On power, the published documents agree on one point, which is that the CDU should be fed from an uninterruptible supply. Whether controllers, actuated valves and the essential air-cooling equipment are backed up as well is decided together with the server shutdown sequence, and the RFQ should state both the requirement and where the boundary of that supply lies.

Three documents that carry an RFQ through to acceptance

Work through the six interfaces and three documents fall out of the exercise: a system schematic with equipment and interface numbers, an interface schedule carrying the duty points and the responsibility split, and a control-and-acceptance matrix running from detection to recovery. Most of the criteria above drop straight into the third one, including flow-balance deviation, supply-temperature stability, the pressure-drop limits on component leak tests and the leak-location points, because they are measurable and recordable at factory acceptance. Site acceptance then answers the different question: how the system behaves once the real cooling source, the real servers and the real monitoring system are connected to it.

ETENZ works from a confirmed server configuration to deliver prefabricated enclosure fabrication and equipment layout, together with the associated electrical, thermal-management, fire-protection and security integration, and can supply the enclosure, prefabricated interfaces or a completed module, including OEM/ODM cooperation, according to the split agreed for the project. Aligning port positions, supports, service clearances, containment and drainage, control points and acceptance documentation before fabrication is what keeps improvised pipework changes and repeated commissioning off the site programme.

To open an enquiry, have the shortlisted server models, the rack schedule, the site conditions and the existing cooling-plant information to hand, plus a note of which equipment each party intends to supply. As the platform and the interface conditions settle, layout, equipment selection and pricing all run off the same inputs. The engineering criteria quoted above come from GB/T 48023-2026, the Chinese national standard for cold-plate liquid cooling systems in data centres, effective 1 February 2027, and from the Chinese association standards T/CECS 1722-2024 and T/CIE 091-2020. All three are recommended documents rather than mandatory ones; outside China they are useful mainly as a well-defined way of phrasing conditions, while what binds a project is the code applicable at its location together with the equipment manufacturers' own limits. Check the clauses and their scope against the original texts.

Tags

liquid cooling interfacesCDU selection criterialiquid cooling RFQrack cooling pressure dropcooling loss response

Share Article

Thank you for sharing ETENZ content