TL;DR

  • Shift from Hardware to Operations: As high-density GPU clusters (100 kW+ per rack) and liquid cooling become standard, physical safety margins shrink, shifting the primary risk of outages from hardware capacity limitations to procedural flaws and human error.
  • Operational Complexity of Liquid Cooling: Direct-to-chip and immersion liquid cooling solve thermal barriers but introduce complex operational demands, such as coolant chemistry, leak detection, and CDU coordination, where minor procedural mistakes carry immediate, expensive consequences.
  • Focus on Long-Term Operational Maturity: AI infrastructure, especially in emerging markets like Central Asia, will ultimately be judged by long-term operational uptime and disciplined management rather than initial construction speed

# # #

As GPU clusters get denser and liquid cooling becomes standard for high-end AI, the decisive variable in AI infrastructure is shifting from how a facility is built to how it is run. Operational standards, not just megawatts, are becoming a critical measure of readiness.

Every conversation about AI infrastructure opens with the same numbers: megawatts secured, racks pushing past 100 kW, GPU counts, gigawatt campuses, interconnect queues. Those numbers matter without power and somewhere to put the heat, nothing scales. But they describe what a facility can theoretically hold, not whether it can carry that load safely at 3 a.m., when a coolant distribution unit alarms and the on-call technician has minutes to decide whether to fail over, shed load, or ride it out. That decision isn’t made by a GPU or a chiller. It’s made by a person following or failing to follow a procedure. And that is where AI data centers are now most exposed.

AI changes the operating model, not just the hardware

Traditional facilities were forgiving: predictable workloads, mature air cooling, settled procedures. AI breaks that. OCP describes AI-driven infrastructure as an order-of-magnitude increase in rack power density, with roadmaps pointing towards 1 MW racks in the next few years. Power volatility adds another challenge: NVIDIA has described rack-level AI workloads swinging from around 30% to 100% utilization and back again in milliseconds, reflecting the synchronized compute and communication phases documented in large-scale AI training research. Loss of coolant flow can quickly translate into thermal throttling, instability or protective shutdowns. At these densities, the coupling between IT load, power and cooling becomes tight and unforgiving. The hard problem is no longer only construction; it’s operational maturity, and the gap between the two is where expensive clusters get lost.

Liquid cooling solves a thermal problem and creates an operational one

Nowhere is this clearer than with liquid. Direct-to-chip liquid cooling is becoming increasingly standard for high-density AI, while immersion remains one of several approaches, but moving water or dielectric fluid through a live data hall introduces operational disciplines that may be unfamiliar to teams accustomed to air-cooled environments: coolant chemistry and conductivity, biological growth and corrosion in the cooling loop, filtration, leak detection and containment, and constant coordination between the facility water system and the rack-level CDU. The OCP Project Deschutes CDU specification targets roughly 2 MW of cooling capacity at 500 GPM with 80 to 90 psi of available pressure, with defined coolant and system requirements for high-density liquid-cooled deployments. At this scale, liquid cooling becomes an operational system in its own right, with consequences for maintenance, monitoring and response procedures. It does not remove the operational problem. It raises the stakes on it.

Human error is now an infrastructure risk

Uptime Institute’s 2025 outage analysis found that 85% of major human-error outages involved either staff failing to follow procedures or flaws in the procedures themselves, and that nearly 40% of organizations had suffered a major outage caused by human error over the previous three years. Its 2026 analysis shows that failure to follow established procedures remains the leading driver of human-error-related outages. Power is still the leading cause of impactful outages overall, so the point is not that procedures replace infrastructure risk. It is that procedures remain one of the clearest controllable weaknesses once the facility is live: whether the runbook exists, whether it is correct, whether anyone follows it, and whether change management catches the conflict before it reaches the floor. In lower-density environments, operational mistakes often had more thermal and electrical margin around them. At 100 kW per rack and above, particularly in liquid-cooled environments, those margins can shrink sharply. The same procedural error can therefore have much faster and more expensive consequences.

What “standards-based operations” actually means

The phrase has to stop being a slogan. It means a facility runs on documented, auditable practice rather than tribal knowledge and real frameworks exist for this. The operations layer has its own independent assessment framework: Uptime Institute’s Management & Operations (M&O) Stamp of Approval evaluates staffing, organization, training, maintenance, operating conditions, planning and coordination independently of the Tier Classification System. A highly resilient design can still be undermined by weak operations. Around it sits the management-system scaffolding the best operators already use: ISO/IEC TS 22237-7:2018, which specifically addresses data centre management and operational processes, ANSI/TIA-942-C for physical infrastructure, and ITIL-style change and incident management that turns change control from a form into a real risk gate. In practice that means clear methods of procedure and emergency operating procedures; change management that accounts for how one modification now ripples across load, cooling, power, and network at once; staff trained for liquid and high density rather than retrofitted from air; telemetry that correlates IT load with thermal and power behavior in real time; and risk management that prices in operational error.

The industry is already moving this way

The clearest signal is OCP’s Open Data Center for AI initiative, launched in October 2025 alongside an industry call for collaboration initiated by Google, Meta and Microsoft and expanded through new workstreams and specifications in 2026. It is explicitly about standardizing not only power and cooling hardware but management telemetry and operational technology and it names the operational problem directly: AI facilities now involve several parties at once (facility owner, hardware owner, workload owner), making service-level agreements and coordination far more complex than the clean colocation SLAs of the air-cooled era. When the largest operators on the planet start writing shared operating specifications, operations not just construction has become a competitive frontier.

Emerging markets: the sharpest version of the problem

This is where I watch it play out most acutely across Central Asia, where AI capacity, power supply, an engineering labor market, and an operating culture are all being built at once. The risk here isn’t only a shortage of megawatts or GPUs; it’s a maturity gap in procedures, people, and change management that never appears on a capex spreadsheet and surfaces only after the first avoidable incident. But the asymmetry cuts the other way too: a greenfield operator carrying no legacy habits can adopt M&O-grade procedures, real change management, and proper telemetry from day one, instead of bolting them on after a decade of “this is how we’ve always done it.” Framed correctly, standards-based operations aren’t bureaucracy that slows an emerging market down. They are the fastest route to closing the reliability gap with mature regions and to skipping mistakes the industry has already paid for.

The real test comes after the ribbon-cutting

AI data centers are becoming national digital infrastructure, and they will be judged not on how fast they broke ground but on whether they run. Power makes deployment possible, cooling keeps density stable, GPUs do the compute but the facility that earns trust over years is the one with correct procedures, trained people, disciplined telemetry, real change control, and risk management that treats human error as seriously as equipment failure. The strongest AI data centers won’t be the ones built first. They’ll be the ones still running predictably while everyone else is firefighting.

# # #

About the Author

Maksim Gavriliuk is an independent digital infrastructure and data center expert, IEEE Senior Member, and DCOS 2026 Review Committee Member. He specializes in data center operations and AI-ready infrastructure, with over 20 years of experience in IT infrastructure projects across Central Asia and Eastern Europe.