Expert

How to Cool a Multi-GPU AI Workstation

Kyle Kyle 04/09/2026 16 min read
Article Tags
Aircooling (6) Compatibility (11) Planning (19) Watercooling (42)
How to Cool a Multi-GPU AI Workstation

Air vs Liquid Cooling, Radiator Capacity and Serviceability

A multi-GPU AI workstation sits in a different thermal class from a gaming PC, and the cooling has to reflect that from the outset. Modern workstation GPUs can draw anywhere from 300W to 600W each, and a high end workstation CPU can add around 350W before memory, storage, VRMs and power supply conversion losses even enter the picture. Air cooling can absolutely work when the GPUs in question are specifically designed for dense multi card deployment and the chassis provides directed airflow. Liquid cooling becomes the more attractive route when acoustics, heat density, GPU spacing or external heat rejection make air impractical. The correct answer here is an engineering decision, not a default assumption carried over from a gaming build.

This guide is about the thermal engineering behind a high density workstation: how much heat you are really planning for, when air is enough, when liquid earns its place, how to size radiators properly, and how to make the system serviceable and monitored over the long term. It assumes a professional workload where uptime and repeatable temperatures matter more than aesthetics.

ComponentRated powerNotes
RTX PRO 6000 Blackwell Max-Q300WDual-slot, designed for up to four-GPU density
RTX PRO 6000 Blackwell Workstation600WDual-slot, high single-GPU performance
Radeon AI PRO R9700300W (TBP)Dual-slot, 12V-2×6, PCIe Gen5
Threadripper PRO 9995WX350W (default TDP)96 cores; needs robust cooling
Example subtotal: 4x 300W GPU + 350W CPU~1,550WBefore memory, VRMs, storage, NIC, fans and PSU losses
A thermal planning illustration from published power ratings. Not measured wall power, not a PSU size, and not a claim that every part peaks at once.

Why AI workstations are thermally different

A gaming PC sees bursty, variable load. An AI or rendering workstation, by contrast, can sustain heavy GPU load, and workload-dependent CPU load, for hours or days at a time, so the cooling system has to handle a sustained heat load rather than a brief peak. On top of that, you are dealing with several high power cards in close proximity, high memory capacity, fast networking and storage, and a requirement to keep temperatures repeatable while the machine remains serviceable. Those priorities (uptime, sustained load, density and fast component replacement) are what should drive every cooling decision from the start.

How much heat can a modern multi-GPU workstation produce?

The scale of the cooling problem is clearest from the published power figures of the components you are likely to be designing around.

NVIDIA RTX PRO 6000 Blackwell Max-Q

The Max-Q version of the RTX PRO 6000 Blackwell is the variant designed for density: 96 GB of GDDR7 ECC memory, a 300W maximum power draw, dual slot active cooling, and explicit support for scaling up to four GPUs for up to 384 GB of combined GPU memory. This is the card NVIDIA intends you to put four of in a workstation, precisely because its power envelope and cooling solution are set up for it.

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

The full Workstation Edition shares the 96 GB GDDR7 ECC memory and 1792 GB/s bandwidth but raises the maximum power to 600W, in a dual slot card with a Double Flow Through thermal design aimed at strong single-GPU performance. One of these represents up to 600W of board power on its own, comparable to the entire peak load of many gaming PCs. That is exactly why you do not simply stack four of them. NVIDIA offers the 300W Max-Q for the four-GPU topology instead. As such, the guidance is straightforward: use the product designed for the density you want.

AMD Radeon AI PRO R9700

On the AMD side, the Radeon AI PRO R9700 offers 32 GB of GDDR6, a 300W total board power, a 12V-2×6 power connector and PCIe Gen5, as an active, dual slot, full height full length card. It sits in the same 300W class as the Max-Q, which is a sensible power point for multi card workstation use.

AMD Threadripper PRO 9995WX

The CPU is not a rounding error here. The Threadripper PRO 9995WX is a 96-core, 192-thread part with up to a 5.4GHz boost, a 350W default TDP and 128 PCIe 5.0 lanes. AMD’s own cooling guidance states that Threadripper and Threadripper PRO CPUs up to 350W require robust cooling, and specifically recommends larger coldplate coverage for parts above 64 cores such as this one. A 350W CPU is a serious cooling load in its own right, not a small addition to the GPU plan.

A thermal planning illustration

It helps to add up the published component power to see the scale, as long as you are clear about what the number is and is not. Four 300W Max-Q GPUs come to 1200W, and with a 350W CPU that is a 1550W subtotal, before you count memory, the motherboard and VRMs, NVMe drives, the NIC, fans and pumps, and power supply conversion losses. Two 600W Workstation Edition cards plus the same CPU reaches the same 1550W subtotal. So a plausible high end CPU and GPU component power budget reaches roughly 1.55 kW before the rest of the workstation is counted.

Be precise about what that figure is. It is a thermal planning illustration built from published maximum and default power ratings. It is not measured wall power, it is not the exact heat produced at every instant, it does not mean every component peaks at once, and it is emphatically not a PSU sizing calculation. Treat it as a sense of scale for the cooling problem, nothing more.

Can you air cool four GPUs?

Yes, in the right hardware. Air cooling is the correct answer when the GPUs are designed for dense multi card placement, the chassis supplies directed, high volume airflow, the inlet air temperature is controlled, the card spacing matches the vendor’s requirements, and the resulting noise and exhaust heat are acceptable. NVIDIA’s 300W Max-Q workstation GPU is direct evidence that dual slot active cards are intentionally engineered for up to four-GPU workstation density on air.

What does not work well is dropping several open air consumer gaming GPUs into a glass gaming pc case, where each card ingests the exhaust of the card below it. Each cooler might be excellent in a single card evaluation and still cook when starved of fresh air and fed its neighbour’s heat. Dense air cooling is about the whole system: purpose built cards plus directed chassis airflow, not just a capable cooler strapped to each GPU.

We do now have useful independent validation for the general principle. A four-GPU workstation using 4x RTX PRO 6000 Blackwell Max-Q cards has been stress tested at a fixed 24°C ambient for three hours, with the GPUs peaking at 85°C while holding the full 300W board power target per card. That does not prove any random four-GPU tower will behave the same way, because the result depended on chassis choice, airflow tuning and thermal monitoring, but it does prove the concept. As such, four 300W class workstation GPUs can be air cooled when the whole platform is designed around them.

ArchitectureWhen it worksAdvantagesMain risks
Dense purpose-built airVendor-designed multi-GPU cards plus directed chassis airflowSimple, serviceableNoise, exhaust heat, high inlet-air demand
Internal liquidLarge chassis with substantial radiatorsCompact GPU blocks, integratedRadiator and case capacity limit
External liquidVery high load, low noise or serviceability priorityHuge heat exchanger, service modularityHose, QDC and power footprint
Manifold / parallel liquidMulti-block professional designsCan reduce aggregate restrictionFlow balance must be validated
Cooling architectures for a high-density workstation. The right one depends on the heat load, the acoustics target and how serviceable the machine must be.

When liquid cooling makes sense

Liquid cooling becomes attractive when it solves a problem air cannot. It lets you remove the bulky air cooler from each GPU, which can improve slot and board clearance depending on the blocks, collect the CPU and GPU heat into one controlled coolant loop, move that heat to large internal or external radiators, and reduce the local airflow you need around each gpu block. For a quiet, high density workstation those are real advantages.

Alphacool ES RTX A4000 Waterblock with Backplate
The Alphacool ES Copper/Carbon water cooler with backplate was developed for the Alphacool Enterprise Series. Due to the positioning of the connections, the hosing of the cooler in the server rack is significantly simplified. The top of the cooler is…
  • SKU: 1023900
  • MPN: 10671
  • EAN: 4250197106719
£169.98£141.65 Inc VatEx Vat
Available
Alphacool ES RTX 6000 Pro Server Edition 1-Slot-Design
Alphacool ES 1-Slot GPU Water Block: Enterprise Solution for RTX Pro 6000 Blackwell The Alphacool ES 1-Slot water block has been specifically developed for professional use in performance-optimized server environments. Thanks to its extremely slim 1-slot design, it enables maximum packing…
  • SKU: 1027359
  • MPN: 19450
  • EAN: 4250197194501
£379.99£316.66 Inc VatEx Vat
Available
Alphacool ES RTX 6000 Pro Workstation / RTX 5090 Founders Edition
Alphacool ES GPU Water Block: Workstation Edition with Active Backplate The Alphacool ES water block has been specifically developed for professional use in performance-optimized workstation and server environments. Thanks to its compact 1.5-slot design and intelligent port positioning, it delivers maximum cooling…
  • SKU: 1028052
  • MPN: 5100182
  • EAN: 4070201001829
  • Available for Collection
£433.62£361.35 Inc VatEx Vat
Available
EK Waterblocks EK-Quantum Vector3 Master RTX 5090 – Plexi
Engineered for the Gigabyte AORUS GeForce RTX™ 5090 MASTER and Gigabyte GeForce RTX™ 5090 GAMING OC 32G models, the EK-Quantum Vector³ Master RTX 5090 – Plexi delivers elite thermal performance for enthusiasts who demand more. Featuring EK’s expanded next-gen cooling engine, pre-cut high-performance…
  • SKU: WAEK-2812
  • MPN: 3831129901148
  • EAN: 3831129901148
£259.49£216.24 Inc VatEx Vat
Available

The important caveat is that a liquid cooled workstation still needs chassis airflow. The VRMs, memory, NIC, storage, motherboard, PSU and anything else you have not put a block on all still rely on moving air. Watercooling the CPU and GPUs does not eliminate the need for case airflow. It changes where the biggest heat loads are handled.

There is also a density threshold where liquid stops being a luxury and starts to look like the cleaner engineering solution. Independent rack systems now exist with up to eight 600W class RTX PRO GPUs, using custom liquid cooling for the GPU, memory and VRM sections, with total cooling capacity in the 6.5 kW range. That is not a normal tower loop, of course, but it illustrates the direction of travel. As heat density rises, moving the heat into a managed liquid loop and a larger heat exchanger becomes far easier to package than trying to brute force the same result with air alone.

How much radiator do you need?

Why rules of thumb are too weak

A crude rule like one 120 mm of radiator per 100W is not good enough for a 1 kW to 2 kW workstation. Radiator capability depends on the radiator design, its frontal area, the fan RPM, the air temperature, the coolant temperature and the resulting coolant to air delta. A rule of thumb ignores all of that.

Design to a coolant to air delta

The professional approach is to choose a maximum acceptable coolant to radiator inlet air delta, then use manufacturer performance data, bench data and your measured heat load to work out the radiator and fan requirement to hold that delta. For very high loads, an external radiator or a rack heat exchanger is often cleaner than trying to cram thick internal radiators into the workstation. Our external radiators guide covers that route in more detail.

Serial versus parallel GPU blocks

With multiple GPU blocks you can plumb them in series, where the coolant passes through each block in turn, or in parallel, where the flow divides between branches. In series, every block sees the full loop flow, the restrictions add up, and the tubing runs are straightforward. In parallel, the total system restriction can be lower, but each branch only receives a share of the flow, and balanced flow depends on the branches having similar restriction. As EK explains, parallel flow divides according to resistance, so a less restrictive path receives more flow.

The key point here is that parallel is not automatically better for many GPUs. It is a hydraulic topology choice, not a performance upgrade. Use blocks and manifolds designed for the arrangement you intend, make sure the loop has adequate total flow, and validate the temperature of each GPU rather than assuming the flow is balanced.

One pump versus two pumps

Alphacool ES Reservoir DDCzero 1U with Pump
£145.99£121.66 Inc VatEx Vat
Available
Add to basket
Alphacool ES Reservoir DDCzero 1U with Pump
The Alphacool ES Reservoir DDCzero 1U with Pump combines a compact reservoir with the powerful Alphacool DDCzero PWM pump. Its low-profile design has been specifically developed for installation in 1U server systems and other water-cooling applications with severely limited installation…
  • SKU: 1025356
  • MPN: 14571
  • EAN: 4250197145718
  • Available for Collection
£145.99£121.66 Inc VatEx Vat
Available
Alphacool ES 2U Reservoir VPP/D5 with flow/temperature sensor
The Alphacool ES 2U Reservoir is a versatile reservoir optimized for use in 2U server cases. The quick-release fasteners on the front ensure quick and easy maintenance. The cooling system can be continuously monitored using the integrated flow and temperature…
  • SKU: 1023975
  • MPN: 13741
  • EAN: 4250197137416
£199.68£166.40 Inc VatEx Vat
Available
Alphacool Eisstation DC80 – AIO Power Unit
The Eisstation DC-LT Solo Top from the Enterprise Solutions series is an extremely compact pump top and is very powerful together with the DC-LT pump.
  • SKU: 1020011
  • MPN: 15388
  • EAN: 4250197153881
£70.68£58.90 Inc VatEx Vat
Available
Alphacool ES 4U Reservoir with D5 Top
£117.95£98.29 Inc VatEx Vat
Available
Add to basket
Alphacool ES 4U Reservoir with D5 Top
The Alphacool ES 4U expansion tank with D5 top is a development for the Enterprise Solutions series. It has an integrated pump top for powerful D5 pumps.
  • SKU: 1021624
  • MPN: 15392
  • EAN: 4250197153928
£117.95£98.29 Inc VatEx Vat
Available

A high density loop can include a CPU block, two to four GPU blocks, QDCs, a flow meter, large or external radiators and long tube runs, which is where a second pump can make sense: for extra head pressure, for redundancy, or to run each pump at a lower duty for a target flow. What you should not do is apply a formula like one D5 per two GPUs. Estimate or measure the loop’s restriction and verify the actual flow. For a critical workstation, adding a second pump for redundancy is a reliability decision, not just a performance one, and one that is best made deliberately.

QDCs, manifolds and serviceability

In a workstation you may need to pull a single GPU, replace a pump or detach the radiator without a full day of draining and rebuilding. Quick release fittings let you define service boundaries: between the chassis and an external radiator, between a GPU manifold and a removable GPU section, or between a test station and a card under test. Our quick disconnect fittings guide covers the details of what is available.

Serviceability has to be planned before any tubing is cut. Ask the awkward questions up front: can one failed GPU be removed on its own, can the pump be replaced, can the radiator subsystem be detached, is there a drain point below the blocks, and can the system be pressure tested in sections? Designing those answers in is what separates a professional build from one that looks impressive but proves miserable to maintain.

Koolance QD3 Quick Disconnect, Straight Female to G1/4 Inch Male Bulkhead – No-Spill, Silver
Koolance patented quick disconnect no-spill coupling with automatic shutoff. Opposite is a panel mountable female G 1/4 BSPP thread for attaching a fitting or barb. The disconnect side will only fit a Koolance QD3 male quick disconnect.
  • SKU: WASA-486
  • MPN: QD3-FTFG4-P
  • EAN: 0829596914993
£22.99£19.16 Inc VatEx Vat
Available
Thermal Grizzly DeltaMate Quick Disconnect Female – F, black
The DeltaMate fittings are water cooling hose connectors for the individual components of a custom water cooling system. Manufactured in the usual Thermal Grizzly quality, the fittings are machined from brass. Brass, as an alloy of copper and zinc, offers…
  • SKU: WAZU-1366
  • MPN: TG-DM-QDC-0002
  • EAN: 4260711993619
£18.34£15.28 Inc VatEx Vat
Available
Thermal Grizzly DeltaMate Quick Disconnect Male – F, Gloss Nickel
The DeltaMate fittings are water cooling hose connectors for the individual components of a custom water cooling system. Manufactured by machine from brass in the usual Thermal Grizzly quality, the fittings benefit from brass being an alloy of copper and…
  • SKU: WAZU-1362
  • MPN: TG-DM-QDC-0320
  • EAN: 4260711993749
£16.50£13.75 Inc VatEx Vat
Available
Thermal Grizzly DeltaMate Quick Disconnect Male – F, black
The DeltaMate fittings are water cooling hose connectors for the individual components of a custom water cooling system. Manufactured in the usual Thermal Grizzly quality, the fittings are machined from brass. Brass, as an alloy of copper and zinc, offers…
  • SKU: WAZU-1368
  • MPN: TG-DM-QDC-0020
  • EAN: 4260711993589
£18.34£15.28 Inc VatEx Vat
Available

Monitoring and alarms

A serious workstation should be instrumented. At the GPU level, monitor core temperature, hotspot or junction where exposed, memory temperature where exposed, board power, clocks and utilisation. For the loop, watch coolant temperature, radiator inlet air, flow and pump RPM, plus fan RPM. At the system level, track CPU package temperature and power, VRM sensors where available, NVMe, the NIC or HBA, and PSU telemetry where the unit provides it. Our GPU hotspot temperature guide covers how to read the GPU side properly. Good monitoring hardware makes all of this considerably easier to implement.

A production oriented machine should alarm on the things that cause damage or downtime: a loss of pump RPM, flow below the expected range, coolant temperature above a threshold, GPU thermal throttling, or a fan bank failure. Set those thresholds from your actual hardware and coolant and tubing limits rather than from any universal number, because the right values depend entirely on the specific components in your loop.

PSU and electrical planning

Do not size the power supply by adding up TDP and board power figures. PSU planning has to account for the vendor’s PSU recommendation, the continuous load, the connector count and type, transient behaviour, cable and current limits, efficiency, the rest of the platform power and some headroom for expansion. For a workstation with a component power budget around 1.5 kW, the electrical supply itself becomes a practical concern: high continuous loads should be considered alongside the circuit, socket and UPS capability. For unusual continuous loads or rack environments, it is better to get qualified electrical advice rather than guessing.

Where does the heat go?

A 1.5 kW class machine under load is effectively a substantial space heater. This is the expectation to set clearly: moving that heat to a giant radiator does not remove it from the room. The heat only leaves the space if the radiator is located in another room, if the heat is ducted or exhausted out, or if the room’s HVAC removes it. An external radiator relocates the heat exchanger. It does not make the heat disappear. Plan for where a continuous kilowatt or more of heat actually ends up, because it has to go somewhere.

A note on data centre accelerators

Bear in mind where heat density is heading, without confusing it with a desktop build. Server accelerators such as NVIDIA’s H200, at up to 700W in the SXM form and up to 600W as an NVL PCIe card, and AMD’s Instinct MI350X, a 1000W air cooled OAM accelerator, show that AI accelerator density is now reaching the 600W to 1000W per accelerator class. Those are server platform parts, not cards you drop into a tower, and they are mentioned here only for context. Keep workstation GPUs, data centre accelerators and consumer cards firmly separate when you plan.

What published testing shows

Because four-GPU workstation hardware is impractical to bench casually, it helps to lean on the independent testing that does exist and then interpret it carefully. The most useful published example to date is a four-GPU air cooled workstation using 4x RTX PRO 6000 Blackwell Max-Q cards and a Threadripper PRO platform, stress tested for three hours at a fixed 24°C ambient. In that configuration, the GPUs reportedly peaked at 85°C while maintaining the full 300W power target per card. I think that is the right way to read the result, not as proof that every tower can do this, but as proof that a properly engineered chassis and airflow path can make four 300W workstation GPUs viable on air.

There is also a helpful lesson in the measured power behaviour of modern AI workloads. Independent inference testing on 300W class workstation GPUs has shown actual power draw varying significantly by model and concurrency, rather than sitting at the board limit constantly. In other words, a 300W rating remains the correct figure for worst case thermal planning, but it should not be mistaken for the exact sustained power draw in every workload. That distinction matters when you interpret temperatures, noise and coolant deltas in the real world.

The other result worth keeping in mind is that multi-GPU only makes sense when the workload actually scales. Independent workstation testing in DaVinci Resolve, for example, shows that GPU effects and some AI workloads can benefit substantially from multiple cards, while other tasks such as Fusion can show limited gains or even awkward behaviour. As such, the cooling solution has to be planned around the workloads that genuinely justify the extra heat, power draw and complexity, not around the idea of multiple GPUs in the abstract.

Frequently asked questions

Can four GPUs run in one workstation?

Yes, when you use GPUs designed for it. NVIDIA’s 300W RTX PRO 6000 Blackwell Max-Q, for example, is positioned for scaling up to four cards for up to 384 GB of combined GPU memory, and independent four-GPU air cooled workstation testing shows that the concept is viable when the chassis and airflow are engineered for it. The key is choosing cards whose power and cooling suit dense placement, and giving them directed chassis airflow, rather than stacking four maximum power or open air consumer cards.

Do AI GPUs need liquid cooling?

Not necessarily. Purpose-built dense workstation GPUs can be air cooled successfully with the right chassis airflow and card spacing, and published four-GPU workstation testing supports that view for 300W class cards. Liquid cooling becomes attractive when you need lower noise, tighter GPU spacing, or to move a large heat load to external radiators. It is a design choice driven by density, acoustics and serviceability, not a universal requirement.

Is a 360mm radiator enough for two GPUs?

It depends on the total heat load, the fans, the coolant to air delta you will accept and whether the CPU is in the same loop. Two 300W GPUs plus a high power CPU is a large sustained load, and a single 360mm radiator may not hold a comfortable delta under continuous use. Design to a target delta using real performance data rather than assuming a radiator size is sufficient.

Should multiple GPU blocks be serial or parallel?

Neither is automatically better. It is a hydraulic choice. Series is simple and gives every block the full flow but adds restriction. Parallel can lower total restriction but splits the flow by branch resistance, so it needs adequate total flow and balanced branches. Use blocks and manifolds designed for your arrangement and validate each GPU’s temperature under load.

Do I need two pumps?

Not automatically. A second pump helps in a very restrictive multi block loop, for extra head, or for redundancy on a critical machine, but many workstation loops run well on one strong pump. Estimate the loop restriction, verify the actual flow, and treat a redundant second pump as a reliability decision rather than a default.


Tagged:
More from Kyle
Jonsbo N6 Review: Did Jonsbo Fix the N5’s Problems? Aug 12, 2026 How Much Does It Cost to Water Cool a PC? Jul 29, 2026 Best AIO Coolers of 2026 Jul 13, 2026

Track My Order

Please only include the numerical part of your order ID, i.e. 177500
Your postcode needs to exactly match what’s on your order, including any spaces
Homepage Modal
Hey There! 👋

 

You're awesome!, let's get you signed up..