A colo contract signed in 2023 for 15kW racks is close to worthless for a 2026 GPU deployment. The building is fine, the power feed might even be adequate, but the room cannot move the heat that a single NVL-class rack throws off. That is the quiet story behind the AI infrastructure buildout: the accelerators arrive liquid-first, and the rooms they land in were designed for air. Buyers who priced their compute around GPU availability are discovering that the chip was never the bottleneck. The cooling was.
This is a decision most technical leaders will face exactly once, under time pressure, with a hardware purchase order already circulating. Get it wrong and you strand a facility or overbuild a water plant you will not fill. Here is how to think about the retrofit before you sign anything.
Air Cooling Ran Out of Headroom, and the Physics Is Not Negotiable
For two decades, data center thermal design was a game of moving more air faster. Hot-aisle containment, higher static-pressure fans, rear-door heat exchangers, each squeezing a bit more density out of the same basic approach. That game is over for AI racks.
Air can practically remove somewhere in the 30-50kW per rack range in a well-run enterprise room. Push into the 60-80kW band and you need rear-door heat exchangers, which are already a form of liquid cooling bolted to the back of the cabinet. Nvidia's GB200 NVL72 rack draws on the order of 120kW, and the roadmap points higher. The Uptime Institute has documented how quickly average rack density climbed once accelerated computing arrived, after years of near-flat growth.
The reason is thermodynamic, not a matter of better engineering. Water carries roughly 3,500 times the heat capacity of air by volume. Once the heat flux at the chip crosses a threshold, no fan array keeps the silicon inside its operating envelope. SemiAnalysis has argued that liquid cooling stopped being a differentiator and became table stakes for frontier training and dense inference clusters. The market data agrees: Dell'Oro Group projects data center liquid cooling revenue growing at a pace that only makes sense if air is being displaced wholesale, not supplemented.
Key takeaway: if your incoming accelerators are from a 2026 generation, air cooling is not a fallback you can rely on while you plan the retrofit. It is a design that no longer applies.
What a Retrofit Actually Touches
"Add liquid cooling" sounds like a rack-level upgrade. It is a facility project. Understanding the layers helps you find where the schedule and the money actually go.
Direct-to-chip (DTC) is the deployable default for 2026. Cold plates sit directly on the GPUs and CPUs, a coolant loop runs through them, and 70-80% of the heat is captured at the source. The remaining heat still needs some air handling for memory, NICs, and power components, so DTC is usually hybrid rather than a full replacement of air. Build's 2026 retrofit analysis walks through what is realistically installable in existing rooms this year, and DTC dominates because it preserves serviceability.
The coolant distribution unit (CDU) is where most retrofit budgets get real. The CDU isolates the clean, controlled coolant loop that touches the chips from the facility water loop that carries heat out of the building. It needs floor space, plumbing, and its own power and redundancy. In many rooms, finding space for CDUs is the first hard constraint.
The facility water loop is the part buyers forget. The CDU has to reject heat somewhere, which means chilled or condenser water piped to the room at the right supply and return temperatures. If your building was never plumbed for this, you are into structural work, permits, and potentially a new cooling plant. ASHRAE's datacom guidance defines the water quality and temperature classes that vendors will hold you to, and the Open Compute Project's cooling-environment specs increasingly define the facility-side interfaces that new hardware assumes.
Immersion cooling deserves a mention because vendors will pitch it. Submerging boards in dielectric fluid removes more heat per rack and can improve efficiency. It also breaks the normal service model, complicates warranties, and demands staff retraining. For a first enterprise retrofit, DTC is the lower-regret choice. Reserve immersion for greenfield builds or when you have a specific density profile that DTC cannot meet.
Key takeaway: the chip vendor sells you cold plates, but your critical path runs through CDU placement and facility water. Scope those first.
Price the Retrofit Against Three Years of Rent
Here is the framework worth internalizing. Before you sign for hardware, model the fully loaded facility retrofit against three years of renting equivalent capacity from a neocloud provider. This mirrors the ownership math we walk through in Rent or Own: GPU-as-a-Service Pricing and Contract Economics, extended to include the thermal plant that the hardware assumes.
The retrofit column includes:
| Cost component | What it covers |
|---|---|
| Cold plates and manifolds | Per-server DTC hardware |
| CDU units and redundancy | N+1 coolant distribution |
| Facility water loop | Plumbing, pumps, possibly new chiller capacity |
| Colo change orders | Contractual permission and space, if applicable |
| Commissioning and staff training | Getting operations ready to run it |
| Downtime and migration | Cost of moving live workloads |
The rental column is comparatively simple: three years of neocloud pricing at your expected utilization, plus data egress and any commitment discounts.
Two variables decide the outcome. Utilization is the first. A retrofit is a fixed cost that amortizes only if the racks run hot most of the time. At steady utilization above roughly 60-70%, owned infrastructure typically wins over three years. Below that, rental's pay-for-what-you-use profile is hard to beat. Certainty is the second. If your workload mix and scale are still moving, a retrofit locks you into a thermal envelope you chose before you understood your demand. Rent until the requirements stop changing.
The counterintuitive part: cooling is frequently the longer lead-time item, not the GPUs. Accelerator allocations can free up in weeks. A facility water loop and CDU installation, with permits and commissioning, can run six months or more. If you plan around chip availability alone, you will have hardware sitting in crates while the room catches up. The power side of this same problem is covered in AI Data Center Power Constraints and Capacity Planning, and cooling deserves the same lead-time respect.
Key takeaway: the three-year rental figure is your break-even reference. If the retrofit cannot beat it at your honest utilization, rent and revisit.
The Specs That Gate Your Schedule
Retrofit projects slip because someone treats the standards as documentation rather than as hard gates. They are gates.
Vendor thermal envelopes are non-negotiable. Each accelerator generation ships with a required coolant supply temperature, flow rate, and pressure drop. Miss the window and the hardware throttles or voids warranty. Get these numbers from the vendor's rack thermal spec before you size the CDU, not after.
OCP cooling-environment specs increasingly define the facility-side interface. As hyperscalers standardize on the Open Compute Project's cooling work, the quick-disconnects, manifold pressures, and temperature classes your hardware expects are converging on these definitions. Building to them keeps you compatible with the next hardware generation instead of designing a one-off loop.
ASHRAE water classes set the quality and temperature bands your loop must hold. Coolant chemistry, filtration, and the difference between a W17 and W32 water class change what plant you need. The ASHRAE datacom series is the reference vendors will cite in a warranty dispute.
Colo CDU and water requirements are the constraint that surprises people. Many existing colocation agreements simply do not permit liquid on the floor, or the facility cannot supply the water loop a CDU needs. Data Center Dynamics has reported on operators scrambling to convert air-cooled halls, and the gating issue is often contractual and structural, not technical. Confirm three things in writing before you commit: that liquid is permitted, that facility water is available at your required temperatures, and that there is physical space for CDUs.
The pattern across all four: the specs turn into schedule risk the moment you discover them late. A room that cannot supply 32C facility water at the required flow is a six-month problem if you learn it during commissioning, and a footnote if you learn it during planning.
Key takeaway: read the vendor thermal envelope, the OCP interface, the ASHRAE water class, and your colo contract before you size anything. Any one of them can stop the project.
Where OpenNash Fits the Retrofit Decision
Most of this decision is infrastructure engineering, and the right answer is often to bring in a mechanical firm for the water plant. Where OpenNash tends to help is upstream of that: the audit and decision framework that determines whether you should retrofit at all, and what workloads justify the fixed cost.
We map the actual compute demand behind the hardware request, because the retrofit-versus-rent math depends entirely on realistic utilization rather than the peak number in a vendor deck. That audit feeds the three-year model directly. On the build side, the automation that makes an owned cluster worth running, orchestration, monitoring, workload scheduling that keeps utilization high enough to justify the capital, is squarely where custom implementation pays off. A cluster that sits at 30% utilization is a rental that would have been cheaper.
The honest positioning: if your demand is spiky or uncertain, we will tell you to rent and revisit, because a stranded water plant is a worse outcome than a neocloud invoice. Ownership makes sense when utilization is steady, the workloads are known, and auditability and control matter to your business. That is the same buildout picture we lay out in The AI Compute Buildout: Players and Dependencies.
Your Next Step
Pull the thermal spec for the exact accelerator SKU on your purchase order and find the required coolant supply temperature and flow rate. Then call your colo or facilities team and ask one question: can the room deliver a facility water loop at that temperature, and is liquid permitted on the floor. If the answer is no or unclear, you have found your critical path, and it is not the GPU. Fix the schedule around that answer before the hardware ships.