GPU power keeps climbing with every generation of AI servers. In dense 8-GPU systems, air cooling is running out of headroom, and more of the heat now has to be handed off to liquid right at the GPU package. That makes the cold plate one of the parts that decides whether the whole server runs cool or not. Pick the wrong one, and you’re looking at thermal throttling at best. At worst, the project goes back to new tooling and a fresh round of validation.
Choosing a server GPU cold plate isn’t the same as choosing a CPU cold plate. A CPU cold plate is built mostly around the socket and retention hardware. A GPU cold plate has to account for the package form factor, several GPUs sharing one baseboard, whether HBM and power components get covered too, and how coolant is split across multiple GPUs. This guide assumes you’ve already settled on direct-to-chip (D2C) cooling. If you’re still weighing D2C against immersion, start with our D2C vs. immersion cooling comparison.
What follows is a selection guide to server and data center GPU cold plates for server OEMs, system integrators, and data center engineering teams, organized around seven key factors.
What Factors Matter When Choosing a Server GPU Cold Plate?
To choose a server GPU cold plate, check these seven factors in order:
- GPU package and server platform
- Real power and heat flux
- Coverage beyond the GPU (HBM, VRMs, switch chips)
- Coolant type and inlet temperature
- Flow distribution and pressure drop across GPUs
- Chassis space, weight, and serviceability
- Leak risk and long-term reliability
1. Start with the GPU Package and Server Platform
The first step in choosing a GPU cold plate is confirming how the GPU sits in the server. The package form factor sets the cold plate’s outline and mounting method, and it decides whether you need one plate per GPU or one assembly for the whole baseboard. Four form factors are common in data centers:
- SXM modules (for example, on NVIDIA HGX baseboards): GPUs plug into a baseboard as modules, usually eight per board. Each cold plate mounts to the module’s mechanical spec, and tubing typically ties the plates on one baseboard into a single assembly.
- OAM modules: OAM is OCP’s mezzanine accelerator module form factor, which differs from a PCIe add-in card, with multiple modules mounted on a Universal Baseboard (UBB). The cold plate logic is similar to SXM, but dimensions and mounting details should come from the OAM spec and the specific module vendor.
- PCIe add-in cards: The GPU is a standalone card, and the cold plate replaces the card’s stock heatsink. Slot spacing and card height force a thin design, and the inlet and outlet have to clear neighboring cards.
- CPU-GPU compute trays (Superchip designs): The CPU and GPU share a board or tray. The cold plate is usually designed for the tray as a whole, with the CPU and GPU on one loop.
Once you’ve confirmed the form factor, gather three things: the module or card mechanical drawings (including keep-out zones), the mounting hole pattern and clamping requirements, and the platform vendor’s reference design, if one exists. APALTEK builds custom liquid cooling modules for 8-GPU server platforms as well as PCIe GPU cold plates, and both kinds of project start from these same documents.
2. Size for Real AI GPU Power, Not the Spec-Sheet Number
Set your GPU cold plate’s thermal targets from the GPU power your system vendor specifies. The same GPU can carry a different power limit depending on the server and cooling method, so the model-level power figures that circulate online often won’t match the machine you’re actually building. Supermicro’s 2-OU (OCP) liquid-cooled HGX B300 system is one example: the company lists eight Blackwell Ultra GPUs at up to 1,100 W TDP each. A cold plate for that system should be designed to that number.
Beyond the power figure, confirm three more things:
- Sustained power, not just rated TDP. AI training often keeps GPUs at full load for long stretches, so the cold plate should be designed for sustained operation.
- Hot spots and heat flux. GPU heat concentrates over the die, so local heat flux there runs far higher than the package average. The HBM stacks around the die have their own heat pattern. The channels in the contact zone should target the real hot spots, not spread flow evenly across the whole package.
- Headroom for the next generation. If the same server is slated for a higher-power GPU later, decide now whether the cold plate should carry extra margin or whether you’ll accept a redesign at the next upgrade.
3. Decide What the Cold Plate Must Cover: GPU, HBM, VRMs, NVSwitch
A server GPU cold plate can cool the GPU alone or extend to the HBM, voltage regulator modules (VRMs), and switch chips like NVSwitch around it, and that coverage choice shapes the plate’s size, weight, and complexity. All of these components run hot, so an early decision is whether liquid cooling handles only the GPUs, leaving the rest to fans, or whether more components go on cold plates as well. On GPU boards, three trade-offs come up.
One-piece vs. split design. A one-piece cold plate uses a single base to cover the GPU and surrounding components. It’s compact and has fewer joints. A split design gives each component its own plate, linked by tubing, which is more flexible to design and easier to replace.
Tolerance stack-up from height differences. GPUs, HBM, and VRMs all sit at different heights, and each has its own tolerance. A one-piece plate has to press evenly on several surfaces at once, which makes base stepping and tolerance control much harder. A common approach is to put thermal interface material (TIM) directly on the GPU contact surface and use compressible thermal pads to take up the height differences everywhere else.
More coverage means fewer fans but a more complex plate. The more components liquid cooling covers, the less the server depends on fans. The cost is a heavier plate, more complex channels, and higher pressure drop. Weigh this against how much air cooling capacity your facility has.
4. Match Thermal Targets to Your Coolant and Inlet Temperature
The same GPU cold plate can perform very differently depending on the coolant and the inlet temperature. That’s why thermal targets have to be set together with loop conditions, not given as a standalone thermal resistance that’s supposed to be “as low as possible.”
Coolant. Single-phase D2C loops typically run deionized water or a propylene glycol-water mix such as PG25. The coolant sets heat transfer performance, and it determines whether the cold plate is compatible with the other wetted materials in the loop.
Inlet temperature. Warm-water cooling, which cuts reliance on chillers, is becoming more common in data centers. According to OCP guidance, single-phase cold plate loops (the TCS) typically run below 49 °C and aren’t expected to exceed 66 °C. The warmer the inlet, the smaller the temperature margin between the coolant and the GPU’s junction temperature limit, and the lower the cold plate’s thermal resistance has to be.
Check against the worst case. Meeting spec under nominal conditions isn’t enough. At the server system level, OCP’s liquid cooling guidelines for OAM systems require that all eight OAMs meet their thermal requirement under the worst case: the lowest coolant flow rate combined with the highest inlet temperature. The same approach works for other GPU platforms. Figure out the minimum flow and maximum inlet temperature your loop can deliver at its worst, then check the cold plate against that point.
Any thermal resistance or temperature target should come with its test conditions. Otherwise, numbers from different suppliers can’t be compared.
5. Plan Flow Distribution Across Multiple GPUs
In a multi-GPU server, every GPU cold plate needs its share of coolant flow, and the combined pressure drop has to fit the loop’s budget. GPU cold plates rarely work alone. In an 8-GPU server, eight cold plates share the in-server tubing, which connects to the rack manifold and the CDU. Together they form what OCP calls the Technology Cooling System (TCS): the loop from the CDU through the manifold and IT equipment and back to the CDU. Cold plate selection has to happen in the context of this loop.
Series or parallel. In a series layout, coolant passes through one plate after another. It’s simple, but downstream GPUs get coolant that upstream plates have already warmed, so they run hotter. In a parallel layout, every plate gets coolant at the same temperature, but flow has to be balanced across the branches. Many real projects combine the two.
Balanced flow. Even a small difference in flow resistance between parallel branches will push more flow toward the path of least resistance. The result is one or two of the eight GPUs running noticeably hotter than the rest. The consistency of each plate’s flow resistance, the tubing lengths, and the connector positions all affect how flow divides.
Pressure drop within budget. The combined pressure drop of the cold plates and in-server tubing has to stay within what the CDU and rack manifold can supply. A cold plate can have excellent thermal resistance, but if its pressure drop blows the budget, real flow falls short of the design point and performance drops with it.
During validation, measure GPU temperatures in the actual upstream and downstream layout, not just on a single plate.
6. Check Chassis Fit, Weight, and Serviceability
A cold plate that hits its thermal targets still has to fit in the chassis and be serviceable.
Height and space. 1U and 2U chassis, as well as OCP’s OU form factor, leave very little height for the cold plate. The plate, retention hardware, tubing, and connectors all have to fit inside that envelope. Tubing also needs to route around memory, NICs, and cables.
Weight. Copper cold plates are heavy to begin with, and wider coverage adds more. The weight of the plate and its retention hardware bears on the GPU modules and baseboard, and shock during shipping and handling amplifies that load. Confirm the baseboard and modules can take it, and add stiffeners if needed.
Quick disconnects and blind mating. For a server to be pulled for service on its own, the cold plate loop needs quick disconnects (QDs). High-density racks are increasingly moving to blind-mate designs. Supermicro’s 2-OU liquid-cooled HGX B300 system for OCP ORV3 racks, for instance, uses blind-mate manifold connections and a modular GPU/CPU tray design. The cold plate’s inlet and outlet positions and connector types have to match how the rack side connects. APALTEK’s GPU liquid cooling modules can be fitted with the QD brand your platform specifies.
7. Plan for Leak Risk and Long-Term Reliability
An 8-GPU server is expensive, and a single leak can cost far more than the cold plate itself. That’s why reliability requirements should be set during selection, not added after prototyping.
Annual failure rate target. OCP’s liquid cooling guidelines for OAM systems state that an annual failure rate of 0.3% or lower is desired for cold plates and coolant loops. A clear target gives the supplier something concrete to design the structure, process, and test plan around.
Where leak detection lives. Beyond the cold plate’s own seal reliability, decide whether leak detection goes inside the server or only at the rack and CDU level, and how the system should respond when a leak is detected.
Integration and shipping. Every step carries risk: installing cold plates in the server, loading servers into the rack, and shipping full racks to the data center. OCP’s integration and logistics white paper recommends that cold plate module vendors supply outgoing quality reports, that cold plates be torqued onto GPUs to the manufacturer’s spec, and that shock and vibration testing account for whether the loop is filled with fluid. If your servers ship as integrated racks, cold plate validation should cover these conditions.
Specific test methods, such as pressure hold, leak testing, and thermal cycling, should be agreed with your supplier up front. APALTEK defines test items based on each customer’s requirements, so spelling them out in your RFQ leads to a more accurate quote and validation plan.
Reference Design or Custom Server GPU Cold Plate?
For many buyers, the real choice isn’t which off-the-shelf cold plate to buy. It’s whether to stick with the server vendor’s reference cold plate or have one custom-built.
A reference design is usually enough when the platform is a standard 8-GPU baseboard, your coverage, loop conditions, and chassis all match the reference design, and speed to market matters most.
A custom cold plate makes sense when:
- The board layout isn’t standard, or you want to cover more components on the same board
- Chassis height, tubing routes, or connector positions differ from the reference design
- Inlet temperature, coolant, or pressure drop budget doesn’t match the reference design’s assumptions
- You want the cold plate matched with your manifold and CDU rather than sourcing each separately and integrating them yourself
APALTEK’s server GPU cold plates and liquid cooling modules are all custom-built, using brazed copper construction. Our AI-8GPU and GPU-OP-8K open-loop modules for 8-GPU AI server platforms, for example, were each developed around platform requirements. Because APALTEK also manufactures manifolds and CDUs, we can match everything from the cold plate to the rack loop and cut down the interface risk that comes with multiple suppliers.
Server GPU Cold Plate Selection Checklist
| Factor | What to Confirm | What Goes Wrong If You Miss It |
| Package form factor | SXM, OAM, PCIe, or CPU-GPU tray; mechanical drawings and mounting spec | Plate doesn’t fit, or uneven clamping causes poor contact |
| Real power | System-vendor GPU power, sustained load, hot spots, upgrade headroom | Throttling at full load, or a full redesign at the next upgrade |
| Coverage | GPU only, or HBM, VRMs, and switch chips; one-piece or split | Surrounding components overheat, or the plate gets too heavy and complex |
| Coolant and inlet temperature | Coolant type, inlet temperature range, worst-case operating point | Passes at nominal conditions, overheats at the worst case |
| Multi-GPU flow | Series or parallel, balanced branch flow, total pressure drop budget | One or two GPUs run hot, or real flow falls short |
| Chassis and service | Height, weight, tubing routes, QDs and blind mating | Difficult assembly, shipping damage, long service downtime |
| Reliability | Annual failure rate target, leak detection level, integration and shipping validation | A leak takes out an entire server |
FAQ
Can one GPU cold plate work across multiple GPU generations? Only if the new GPU’s package size, mounting method, and power all fall within the original cold plate’s design range. GPU generations often change module outline, retention hardware, and power, and when they do, the cold plate usually needs a redesign. If an upgrade is planned, check the margin during selection rather than waiting until the new GPU arrives.
What thermal resistance should a server GPU cold plate have? There’s no universal number. You work it out from your own operating conditions. One common approach, based on junction-to-inlet thermal resistance, is to subtract the coolant inlet temperature from the GPU’s maximum allowable junction temperature, then divide by GPU power. The result is the maximum allowable thermal resistance for the entire heat path, from the die through the package, the TIM, and the cold plate to the coolant. Because the package and TIM take up part of that budget, the cold plate itself has to come in well below that figure. The higher the inlet temperature or the power, the lower the thermal resistance has to be. When comparing suppliers, make sure their numbers were measured under the same test conditions.
How is a server GPU cold plate different from a gaming GPU water block? They work on the same principle but are built for different jobs. A server GPU cold plate is designed around data center GPU module form factors and rack-level loops, and it has to handle sustained full load, flow balancing across multiple GPUs, QD-based servicing, and strict leak reliability requirements. A gaming GPU water block is built for a single card in a PC on a consumer loop, where easy installation and aesthetics carry more weight.
Conclusion
A server GPU cold plate can’t be judged on thermal resistance alone. The GPU package, real power, coverage, loop conditions, multi-GPU flow distribution, chassis space, and reliability requirements together determine how it performs once it’s inside the server.
If you’re developing GPU liquid cooling for an AI server, send APALTEK’s engineering team your GPU platform and package (SXM, OAM, or PCIe), board layout, GPU power and maximum junction temperature, coolant and inlet temperature, available flow and pressure drop budget, and chassis height. We’ll come back with a recommended coverage approach and channel design, run a thermal simulation assessment, and arrange prototypes and testing.