TL;DR
- Shift Focus from Deployment to Long-Term Operations: While the urgency to bring AI capacity online drives rapid installation, long-term success depends on treating liquid cooling as an ongoing operational discipline rather than just a quick procurement project.
- Holistic Resilience & Blast Radius Control: Evaluating system resilience requires looking beyond redundant Cooling Distribution Units (CDUs) to minimize the overall failure “blast radius” across the entire fluid distribution and control network.
- Proactive Performance Management & Adaptability: Leveraging real-time operational data allows operators to catch performance drift before component failure occurs, ensuring cooling systems continuously adapt as AI hardware and dynamic workloads evolve.
# # #
by Kevin Roof, Director of Offer & Capture Management, LiquidStack
Much of today’s discussion around liquid cooling centers on deployment. Which CDU should I buy? Which cooling technology should I choose? What rack density can I support? These are important questions, but they represent only the beginning of the liquid cooling journey.
There is an understandable urgency to bring AI capacity online as quickly as possible. But in some cases, the pace of deployment is beginning to outstrip the operational maturity surrounding liquid cooling. Many organizations are implementing these systems for the first time, often without the decades of established practices that exist around more traditional data center infrastructure.
As AI infrastructure grows more critical the question therefore needs to shift from “How do we deploy liquid cooling?” to “How do we keep these systems operating efficiently for the next decade?”
Installation is measured in weeks. Operations are measured in years.
Liquid Cooling Operations Are an Engineering Discipline
Many operators still approach liquid cooling primarily as a procurement and deployment project: find the right CDU, integrate it into the environment, bring the GPU clusters online and move on to the next capacity challenge. But liquid cooling systems are infrastructure platforms requiring careful and continuous management.
Design establishes the conditions for long-term performance. Commissioning validates that those decisions work in practice. Operations determine whether that performance can be sustained as workloads, hardware and operating conditions change.
That last point is increasingly important. Traditional data center operations were largely built around maintaining equipment. With liquid cooling, the technology cooling system, facility infrastructure, controls and IT equipment operate as a highly interconnected system. Changes to rack configurations, accelerators, control strategies or even routine maintenance can affect behavior elsewhere in the cooling loop.
That is why liquid cooling operations increasingly need to be viewed as an engineering discipline rather than simply a maintenance function.
Fluid Management Means More Than Filling the Loop
One of the biggest misconceptions surrounding liquid cooling is that coolant is simply another utility, much like chilled water flowing through a facility.
In reality, fluid management is one of the most important determinants of long-term system reliability. The fluids circulating through a liquid cooling system must remain compatible with every material they encounter, from piping and seals to cold plates and heat exchangers.
Maintaining proper chemical balance helps minimize corrosion, prevent scaling, reduce biological growth and preserve consistent thermal performance. Small deviations may not create immediate problems, but over time they can gradually sap efficiency and shorten the lifespan of critical components. Internal channels on advanced coldplates can be as small as 100µm, which is about the diameter of a human hair. Even minor biological growth or debris from scaling can easily clog coldplate channels or the filters in place to protect them.
The important mindset shift, then, is to stop treating fluid quality as something checked only during scheduled service. It is an ongoing operational condition that can affect virtually every component in the system.
Resilience Is About More Than Redundant CDUs
High availability has always been a cornerstone of data center design, but redundancy becomes even more important as AI workloads grow in value and computational intensity.
Most operators understand concepts such as N, N+1 or N+2 configurations. Similar principles apply to liquid cooling, but redundancy should extend well beyond the CDU itself.
One mistake is evaluating resilience too narrowly at the component level. A redundant CDU does not necessarily create a resilient cooling architecture if another shared element can interrupt cooling across an entire compute pod.
A better way to think about resilience is in terms of blast radius: if something does fail, how much of the environment does that failure affect?
A resilient system considers pumps, heat exchangers, sensors, valves, controls, the broader fluid distribution network, as well as the upstream power architecture and TCS loop layout. The goal is not simply to eliminate individual points of failure. It is to contain failures and ensure that maintenance or component replacement can occur without unnecessarily disrupting compute operations.
Maintenance Is Becoming Performance Management
Liquid cooling also creates an opportunity that operators should take advantage of: these systems generate a wealth of operational data.
Flow rates, pressure differentials, supply and return temperatures, pump performance, fluid quality and leak detection collectively provide a detailed view into overall system health. The question is whether organizations use that information simply to react to alarms or to understand how system performance is changing over time.
Preventive maintenance has traditionally centered on fixed intervals. Fluid is sampled, filters are inspected, pumps are verified, sensors are calibrated and valves are tested according to a schedule. Those activities remain essential, but they no longer tell the whole story.
In increasingly dynamic AI environments, understanding how the system behaves between maintenance intervals can be just as important.
Small deviations in pressure, pump current, temperature or flow stability may not indicate an immediate problem. But they can provide early evidence that conditions are drifting. By the time a component fails, the system may have been signaling deterioration for weeks.
That turns maintenance into performance management. Instead of waiting for failures, operators can identify trends, investigate abnormalities and schedule service before small changes become production issues.
There is an organizational dimension here as well. Liquid cooling increasingly blurs the boundaries between facilities, IT teams, tenants and equipment vendors. If different stakeholders can make changes that influence flow, pressure or thermal behavior, organizations need clear ownership around who monitors the system, who can change parameters and how visibility is shared.
Operational resilience increasingly depends on coordination as much as hardware.
The System You Install Is Not the System You Will Operate
Perhaps the biggest mistake organizations can make is designing and operating exclusively for today’s realities.
AI infrastructure is evolving at an extraordinary pace. Cooling systems deployed today may be expected to support several generations of accelerator technology, each bringing different thermal characteristics, flow requirements and operating behaviors. And it is not simply a question of supporting more heat.
Training environments can introduce rapid swings in compute demand, requiring cooling systems to respond dynamically. Inference environments may create steadier, sustained loads. Hardware refreshes can change hydraulic requirements even when the underlying cooling infrastructure remains physically unchanged.
That means successful commissioning does not establish a permanent operating state.
When accelerators change, racks are reconfigured or significant maintenance alters the cooling loop, operators should reassess whether the assumptions under which the system was originally commissioned still apply. In some cases, changes may warrant new control strategies, different setpoints or even renewed system validation.
A well-operated liquid cooling system should not be defined by its ability to remain unchanged. It should be defined by its ability to adapt while continuing to deliver predictable thermal performance.
Installation Is Only the Beginning
Liquid cooling has become one of the foundational technologies enabling the AI era, but technology alone does not guarantee success.
The true measure of a deployment is not how quickly the equipment is installed or how well the system performs on commissioning day. It is whether that infrastructure continues to deliver reliable, efficient cooling years later as workloads expand, hardware changes and operating conditions evolve.
That requires a different mindset. Installation is not the finish line, and commissioning does not freeze a cooling system in time. Liquid cooling must be continuously observed, maintained, adjusted and, when conditions materially change, revalidated.
The industry has spent enormous energy learning how to deploy liquid cooling for AI. Now we need to become equally good at operating it.
# # #
About the Author
Kevin Roof is the Global Director of Offer & Capture Management for LiquidStack. A mechanical engineer and PMP with over a decade of experience in data center cooling, Kevin brings invaluable insights and thought leadership to the liquid cooling space.