
How to design, deploy, and scale liquid-cooled infrastructure for high-density AI workloads
Modern AI workloads are driving unprecedented rack densities and power draws, pushing traditional air-cooled data centers to their limits. For example, standard racks have historically had an average installed power of 10-20kW, but the current power requirements for a full rack of NVIDIA’s latest Blackwell NVL72 reaches 120kW. The next generation of AI compute could push the power demand even further to 400kW and beyond.
As AI and high performance computing (HPC) continue to increase in scale and complexity, liquid cooling has become a necessity. The direct-to-chip liquid cooling approach in particular is becoming a mainstream solution for modern data centers because it enables highly localized and efficient heat dissipation for high-density AI infrastructure.
In this playbook, we’ll discuss how to navigate the key decisions around AI infrastructure, including power, cooling, facility constraints, architecture designs, and more. Read on to learn our best practices for adapting data centers for AI workloads with liquid cooling.
Clarify AI Strategy and Use Cases
Start by anchoring infrastructure decisions to specific business use cases because different AI initiatives will require very different hardware and architecture designs. For example, large language model (LLM) training workloads typically run on GPU-dense clusters with massive power draws, while real-time inferencing requires efficiency and low latency using high-capacity CPUs and GPUs.
This means it’s important to understand which AI workloads are in-scope, and what their requirements are in terms of compute, latency, and throughput. Then you can set a realistic target for compute density over a 3-5 year horizon to avoid over- or under-building infrastructure.
Get a Baseline of Current Data Center Estate
Before designing new AI infrastructure, you’ll want to get an understanding of your existing environment and constraints. You should quantify your actual power envelope at the facility level, including available capacity and future expansion potential. Similarly, you should evaluate your existing cooling plant and your maximum achievable rack density on air.
Space, floor loading, and water availability may introduce additional constraints on upgrading power capacity and the feasibility of deploying liquid cooling systems. By understanding the limitations of your existing data center facility, you can identify where upgrades or redesigns may be required before introducing high-density AI infrastructure.
Define AI Factory Design Principles
Now you can translate your strategy and constraints into clear principles centered around modularity, scalability, efficiency, and risk tolerance. We suggest designing vendor-agnostic AI pods that can be easily adapted to different use cases and scaled incrementally as AI technologies mature.
We also recommend defining an AI Factory approach for moving workloads or use cases along a standardized production line that includes infrastructure, data, AI platforms, and an operating model. This ensures AI initiatives can move more quickly and reliably from exploration to production. While building your AI Factory roadmap, it’s helpful to decide whether you’re optimizing for a few ultra‑dense AI pods or moderate density across your broader data center estate.
Select a Liquid Cooling Strategy and Reference Architecture
Now you can choose which liquid cooling approach fits your business strategy and data center constraints:
- Direct-to-chip cooling: Uses liquid coolant applied directly to cold plates attached to CPUs and GPUs. This can support extremely dense server configurations and is ideal for advanced AI processing, machine learning, and large-scale analytics workloads.
- Rear-door heat exchangers: Replaces standard rack doors with liquid-cooled heat exchangers. This is a transitional solution that can enhance the efficiency of air-cooled environments by removing heat at the rack level.
- Immersion cooling: Submerges entire servers into non-conductive fluid for more uniform heat dissipation at the system level. This requires greater infrastructure investment because it’s more complex to retrofit into traditional data centers.
In addition, a hybrid approach combining liquid and air cooling could make sense if it’s not possible to fully retrofit your entire data center. This enables you to strategically deploy liquid cooling for high-density AI racks while maintaining existing air-cooled infrastructure for less demanding workloads.
Next you can map vendor and OEM reference designs to your mechanical, thermal, and electrical requirements. Leading vendors like CoolIT, Motivair, and Vertiv offer liquid cooling technologies and supporting infrastructure with different capabilities and price points that are worth evaluating.
Engineer the Power Architecture for AI-Class Racks
The fundamental constraint for AI infrastructure today is power consumption, so it’s crucial to ensure there’s enough capacity available at both rack and facility levels. This means creating engineering power budgets per rack and planning branch circuit designs with appropriate redundancy tiers.
Besides powering the racks themselves, the power budget needs to incorporate the additional load requirements of chillers, pumps, and auxiliary cooling systems for liquid cooling as well. It’s also important to consider upstream power distribution, and design redundancy along with a metering strategy to avoid stranded capacity.
Design the Cooling Architecture and Fluid Networks
Now you can translate your cooling strategy into a concrete plant and plumbing design that’s reliable and realistic based on your data center constraints. This includes integrating secondary fluid loops for liquid cooling systems with the primary fluid loop of the facility infrastructure like chilled water systems and coolant distribution units (CDUs).
