What Is Liquid Cooling? How It Works and Why AI Data Centers Need It

[Global] Success Blueprints|2026. 9. 1. 07:11
반응형

Artificial intelligence is transforming the data center industry, and the investment story is expanding far beyond GPUs.

For the first phase of the AI infrastructure boom, investors focused primarily on semiconductor performance. NVIDIA GPUs, high-bandwidth memory, advanced networking, and leading-edge semiconductor manufacturing became some of the most visible beneficiaries of surging AI capital expenditures.

But as AI clusters become larger and more power-intensive, another constraint is becoming increasingly important

Heat.

A high-performance GPU is valuable only if it can operate reliably at scale. As more GPUs are packed into increasingly dense server racks, the challenge is no longer simply obtaining enough compute.

Data centers must also supply enough power and remove the enormous amount of heat that compute generates.

That is why liquid cooling is emerging as a critical part of the AI infrastructure stack.

Liquid cooling technology removing heat from high-performance CPUs and GPUs in an AI data center
An overview of liquid cooling technology and how coolant removes heat directly from high-performance CPUs and GPUs in AI data centers.

What Is Liquid Cooling?

Liquid cooling is a thermal-management technology that uses water or specialized coolant to absorb and transfer heat away from high-temperature components such as CPUs and GPUs.

Traditional data centers have relied heavily on air cooling.

Fans move hot air away from servers while HVAC and cooling systems circulate conditioned air through the facility.

That approach remains practical for many conventional workloads.

AI servers, however, are changing the thermal equation.

Modern AI systems concentrate powerful GPUs, high-bandwidth memory, networking equipment, and other components within a relatively small physical footprint.

The result is significantly greater heat generation per rack.

Liquid can transport heat much more effectively than air, allowing cooling systems to remove heat closer to where it is actually generated.

A simple way to think about the difference is this

Air cooling is like cooling an entire hot room with air conditioning. Liquid cooling is closer to attaching the cooling system directly to the equipment producing the heat.

 

Why AI Is Driving Demand for Liquid Cooling

Rising GPU power density and heat generation increasing liquid cooling demand in AI data centers
An illustration of how larger AI models and increasingly powerful GPUs drive higher rack power density, heat generation, and demand for advanced cooling systems.

Liquid cooling itself is not a new technology.

What has changed is the architecture of computing.

The AI infrastructure cycle creates a straightforward chain reaction

Larger AI models → More GPU computation → Higher electricity consumption → More heat → Greater cooling requirements

One of the most important concepts for investors to understand is power density.

Power density measures how much electricity computing equipment consumes within a limited physical space, particularly inside an individual server rack.

As GPU performance increases and more accelerators are installed within each rack, rack-level power requirements can rise dramatically.

NVIDIA's GB200 NVL72 rack-scale AI system, for example, operates at roughly the 120-kilowatt level and incorporates liquid cooling into its architecture.

This illustrates an important shift in AI infrastructure economics.

The question is no longer simply

How many GPUs can a data center buy?

Increasingly, the more important question is

How many GPUs can the facility power, cool, and operate reliably?

That distinction matters because installed GPUs that cannot run efficiently do not translate into usable computing capacity.

 

How Does Liquid Cooling Work?

Direct-to-Chip, immersion cooling, rear door heat exchanger, and CDU systems for AI data center cooling
A comparison of Direct-to-Chip cooling, immersion cooling, and rear door heat exchangers, including the role of CDUs in transferring heat away from AI servers.

There are several approaches to liquid cooling, but three architectures are particularly relevant to the data center industry

Direct-to-Chip cooling, Immersion Cooling, and Rear Door Heat Exchangers.

Direct-to-Chip Cooling

Direct-to-Chip, or DTC, cooling brings liquid directly to the major heat-generating components inside a server.

A metal cold plate is installed on components such as CPUs or GPUs. Coolant circulates through the cold plate, absorbs heat, and transports it away from the processor.

The basic thermal path looks like this

GPU/CPU → Cold Plate → Coolant → CDU → Heat Exchanger → External Heat Rejection

One of the most important components in this architecture is the Coolant Distribution Unit, or CDU.

The CDU manages the interface between the server-side cooling loop and the facility-side cooling infrastructure.

It helps control coolant temperature, pressure, and flow while transferring heat from the IT equipment to the broader data center cooling system.

This means liquid cooling is not simply a matter of attaching water lines to GPUs.

It requires an entire thermal-management ecosystem involving

Cold plates → Manifolds → Quick disconnects → Pumps → CDUs → Heat exchangers → Monitoring systems

That ecosystem is where the investment opportunity begins to broaden.

 

What Is Immersion Cooling?

Immersion cooling takes a fundamentally different approach.

Instead of circulating coolant through cold plates, servers or computing equipment are submerged directly in a specially designed dielectric fluid.

This fluid does not conduct electricity like ordinary water.

Because the server components come into direct contact with the cooling medium, immersion systems can remove significant amounts of heat and potentially reduce reliance on internal server fans.

The technology can be attractive for extremely dense computing environments.

But thermal performance is only part of the equation.

Immersion cooling may require changes in server architecture, maintenance procedures, fluid management, component compatibility, and operational standards.

For investors, the relevant question is therefore not simply which technology removes the most heat.

It is

Which cooling architecture can be deployed reliably, economically, and at scale?

 

Rear Door Heat Exchangers

A Rear Door Heat Exchanger provides another approach, particularly for facilities attempting to increase rack density without completely redesigning existing server infrastructure.

A liquid-cooled heat exchanger is installed behind the server rack.

Hot exhaust air leaving the servers passes through the exchanger, where heat is transferred into the liquid cooling loop.

Unlike Direct-to-Chip cooling, liquid does not travel directly to the GPU or CPU.

This can make rear-door systems useful in certain retrofit environments where operators want to retain much of their existing air-cooled architecture while improving thermal capacity.

The broader takeaway is important

There may not be one universal winner.

Different data centers can adopt different cooling architectures depending on rack density, existing infrastructure, capital costs, maintenance requirements, and workload characteristics.

 

Why CDUs Matter

The CDU deserves particular attention because it functions almost like the heart of many liquid-cooling architectures.

It manages the thermal connection between high-value computing hardware and facility infrastructure.

As liquid cooling adoption expands, demand therefore does not necessarily accrue to a single category of cooling equipment.

An entire supporting ecosystem can grow around it

Cold plates, pumps, valves, manifolds, connectors, CDUs, heat exchangers, sensors, monitoring software, and integrated thermal-management systems.

This is an important distinction for investors looking beyond headline AI semiconductor names.

The opportunity may exist across the infrastructure required to make increasingly powerful computing systems physically operable.

 

Liquid Cooling and Data Center Efficiency

Another important data center metric is Power Usage Effectiveness, or PUE.

PUE is calculated as

PUE = Total Data Center Energy Consumption ÷ IT Equipment Energy Consumption

The closer PUE approaches 1.0, the greater the share of facility power being used directly by computing equipment rather than supporting infrastructure.

Liquid cooling can potentially reduce some of the energy required for thermal management.

But investors should be careful not to assume that liquid cooling automatically produces the same efficiency gains at every facility.

Climate, utilization rates, cooling architecture, electrical equipment efficiency, facility design, and operating conditions can all influence the outcome.

PUE also does not tell the entire story.

A facility can improve its infrastructure efficiency while still underutilizing expensive GPUs.

For AI data centers, the more economically meaningful question may ultimately be

How much useful AI computation can the facility produce from each unit of available power?

 

The Bigger Opportunity: Compute Density

Liquid cooling increasing GPU density and usable AI compute capacity in high-density data centers
An illustration showing how thermal capacity can become a major AI infrastructure bottleneck and how better cooling enables higher GPU density and more usable compute.

This brings us to the real economic significance of liquid cooling.

The biggest advantage may not simply be lower electricity costs.

It may be higher compute density.

Every data center operates under physical constraints.

There is only so much space.

There is only so much electricity available.

And there is only so much heat the facility can remove.

If thermal capacity becomes the limiting factor, buying additional GPUs does not solve the problem.

They may simply be impossible to deploy within the existing facility.

But if cooling capacity improves, operators can potentially install more computing resources within the same footprint.

The relationship becomes

Better cooling → Higher GPU density per rack → Greater compute density

That is why thermal management is becoming economically important to the AI infrastructure cycle.

The competitive advantage of an AI data center is increasingly determined not only by the GPUs it owns, but by how densely and reliably those GPUs can operate.

 

AI's Power Problem Makes Cooling More Valuable

Cooling also cannot be separated from the broader electricity problem facing AI infrastructure.

The International Energy Agency expects global data center electricity consumption to reach roughly 945 TWh by 2030, approximately double the 2024 level.

AI-optimized accelerated servers are expected to be an important contributor to that growth.

The challenge is that data centers can sometimes be constructed faster than new generation, transmission, and distribution infrastructure.

That creates two basic options for operators

Secure more electricity.

Or use existing electricity more efficiently.

Liquid cooling can contribute to the second strategy.

If less facility power is required to manage heat, more of the available electrical capacity can potentially be directed toward actual computing workloads.

This turns thermal management from a facility expense into a strategic infrastructure consideration.

 

Where Could AI Infrastructure Spending Move Next?

AI infrastructure investment shifting from GPUs to liquid cooling, power systems, CDUs, and data center infrastructure
A visualization of AI infrastructure spending expanding from GPUs toward power and liquid cooling, including cold plates, pumps, CDUs, heat exchangers, monitoring systems, and data center infrastructure.

The early AI investment cycle was dominated by GPUs.

But a functioning AI cluster requires far more than accelerators.

GPUs require HBM.

Large clusters require high-speed networking.

Servers require electricity.

And all that electricity ultimately becomes heat that must be removed.

The AI infrastructure spending chain can therefore expand

GPU → HBM → Networking → Power Infrastructure → Cooling Infrastructure → Data Center Facilities

This is one of the most important concepts for long-term AI investors.

Capital does not only chase the most obvious growth technology.

It also moves toward the bottlenecks created by that growth.

When GPU supply improves, the next constraint may appear elsewhere.

It could be memory.

It could be networking.

It could be electricity.

And increasingly, it could be cooling.

 

The Liquid Cooling Value Chain

If liquid cooling adoption expands, the beneficiaries could extend well beyond traditional cooling-equipment manufacturers.

Potential areas of the ecosystem include

Cold plates that transfer heat directly from processors.

CDUs that distribute and control coolant.

Pumps, valves, piping, manifolds, and quick-disconnect systems that create reliable cooling loops.

Heat exchangers and facility cooling equipment that reject heat outside the computing environment.

Dielectric fluids used in certain immersion-cooling systems.

Sensors and monitoring software that detect leaks and optimize thermal performance.

Data center engineering and construction firms capable of integrating electrical and cooling infrastructure.

And ultimately, data center operators whose facilities can support high-density AI racks.

The key competitive advantage may not be any individual component.

It may be system reliability.

When thousands of expensive accelerators operate together, even a relatively small cooling failure can create significant operational and financial consequences.

Companies capable of combining thermal performance with reliability, monitoring, maintenance, and data center integration may therefore have stronger competitive positioning than businesses selling commodity cooling hardware alone.

 

The Risks Investors Should Watch

Liquid cooling is not a guaranteed winner for every company exposed to the theme.

Several risks deserve attention.

First is capital intensity.

Retrofitting existing facilities for Direct-to-Chip cooling may require new piping, CDUs, heat exchangers, and modifications to facility infrastructure.

Second is reliability and leakage risk.

Bringing liquid close to high-value electronics increases the importance of connectors, monitoring systems, and maintenance procedures.

Third is technology standardization.

Direct-to-Chip, immersion cooling, rear-door systems, and hybrid approaches may coexist rather than converge immediately on a single architecture.

Fourth is profitability.

A rapidly growing market attracts competitors.

Industry revenue can rise while equipment pricing and margins decline.

Investors therefore need to distinguish between growth in the overall cooling market and sustainable economics for individual companies.

 

Will Liquid Cooling Replace Air Cooling?

Probably not everywhere.

Conventional web servers, storage systems, and other lower-density workloads may continue to use air cooling because it remains economical and operationally familiar.

High-performance AI training clusters are different.

As rack density rises, liquid cooling becomes increasingly compelling.

The future data center may therefore be less about

Air cooling versus liquid cooling

and more about

Air cooling for conventional workloads + liquid cooling for high-density AI infrastructure.

Technology transitions rarely require the old architecture to disappear completely.

What matters economically is identifying the point where the new technology becomes superior for a specific workload.

 

What Long-Term Investors Should Watch

The liquid-cooling investment thesis should not begin with the question

"Which liquid cooling stock will go up?"

A better framework is to follow the movement of infrastructure bottlenecks.

The AI cycle can evolve from

Compute → Memory → Networking → Power → Cooling

Investors should watch several variables.

Is rack-level power density continuing to rise?

Which cooling architecture is gaining adoption in new versus existing facilities?

Which suppliers are being qualified by major server manufacturers, hyperscalers, and data center operators?

Are companies generating improving cash flow and margins, or merely benefiting from temporary AI-related order growth?

And perhaps most importantly

Is the company selling a one-time piece of equipment, or providing infrastructure, maintenance, technology, or services that remain necessary as long as AI data centers continue operating?

That distinction can become increasingly important as the AI capital expenditure cycle matures.

 

Final Takeaway

Liquid cooling is becoming an important part of AI infrastructure because the economics of computing are changing.

The chain increasingly looks like this

GPU performance → Electricity → Heat → Cooling → Usable Compute

The fastest accelerator in the world has limited economic value if a data center cannot supply enough electricity or remove enough heat to operate it reliably.

That means investors analyzing AI infrastructure should increasingly look beyond GPU shipment numbers.

Rack-level power density, Direct-to-Chip adoption, CDUs, electrical capacity, cooling CAPEX, facility retrofits, and high-density data center design are becoming part of the same investment story.

The next major AI bottleneck does not necessarily have to occur inside a semiconductor.

As computing performance increases, the physical infrastructure required to power and cool that performance becomes increasingly valuable.

For long-term investors, the opportunity may therefore be less about chasing the hottest technology and more about identifying the companies solving the bottlenecks that the technology itself creates—and doing so with durable cash flow and sustainable economics.

반응형

댓글()