DEV Community

Da
Da

Posted on • Originally published at sensaka.com

How to Improve PUE Without Compromising Reliability

Power Usage Effectiveness is one of the most widely used data center efficiency metrics.

It compares total facility energy with the energy consumed by IT equipment. The closer the result is to 1.0, the smaller the share of energy used by cooling, power conversion, lighting, and other supporting systems.

This makes PUE useful, but it can also create the wrong behavior.

A team that pursues a lower number without considering resilience may reduce safety margins, disable redundancy, raise temperatures too aggressively, or make changes that improve the metric while increasing operational risk.

The objective should not be the lowest possible PUE at any cost.

The objective should be to reduce avoidable facility energy while continuing to protect equipment, workloads, and business services.

Start with a consistent measurement boundary

PUE improvement begins with reliable measurement.

The basic relationship is:

Total facility energy divided by IT equipment energy.

The calculation appears simple, but results can vary according to where and how energy is measured.

Questions to define include:

  • Which facility loads are included?
  • Where is total energy measured?
  • Where is IT energy measured?
  • Are shared building systems included?
  • Are office areas excluded?
  • Is generator fuel included?
  • Are cooling pumps included?
  • Are lighting and security included?
  • Is measurement continuous or based on samples?
  • Are annual and peak values reported separately?

If the boundary changes, the result may improve without any real efficiency gain.

For example, moving a cooling load outside the measurement boundary does not reduce energy consumption. It only changes the accounting.

The guide to what PUE is provides the core definition and explains the main factors that influence the metric.

Use annual PUE, not only a single snapshot

PUE changes with:

  • Weather
  • IT load
  • Cooling mode
  • Equipment utilization
  • Time of day
  • Maintenance activity
  • Seasonal conditions
  • Facility occupancy

A value measured during a cool night at high IT load may look excellent. The same facility may perform very differently during a hot afternoon or at low utilization.

Useful reporting should include:

  • Annual PUE
  • Monthly PUE
  • Daily trend
  • Peak conditions
  • Low load conditions
  • Seasonal variation
  • Load level
  • Weather context
  • Planned maintenance periods
  • Meter quality

An interactive PUE calculator can help teams understand the relationship between facility energy and IT energy, but operational improvement depends on continuous, consistent data.

Improve the denominator carefully

PUE can improve when IT load increases, even if total energy also increases.

This happens because fixed facility loads are spread across more IT consumption.

For example, lighting, security, controls, and some cooling systems may consume a similar amount of energy whether the facility is lightly or heavily loaded.

This means a lower PUE can result from:

  • Better facility efficiency
  • Higher IT utilization
  • More IT equipment
  • Changes in measurement
  • Seasonal weather

These are not equivalent.

Teams should therefore report PUE alongside:

  • Total facility energy
  • IT energy
  • IT utilization
  • Rack utilization
  • Workload output
  • Cooling energy
  • Power losses
  • Electricity cost

A lower PUE should reflect real efficiency improvement rather than simply a larger denominator.

Airflow management is often the safest first step

Many facilities cool more aggressively than necessary because cold air and hot air are not controlled effectively.

Common airflow problems include:

  • Missing blanking panels
  • Unsealed cable openings
  • Open rack spaces
  • Poor containment
  • Obstructed floor tiles
  • Incorrect perforated tile placement
  • Hot air recirculation
  • Cold air bypass
  • Uneven rack density
  • Poor equipment orientation
  • Excess airflow

These problems cause the cooling system to work harder while some equipment still receives air at the wrong temperature.

Low risk improvements include:

  • Install blanking panels
  • Seal cable openings
  • Remove airflow obstructions
  • Align equipment intake and exhaust direction
  • Balance perforated tile placement
  • Close unused rack openings
  • Separate hot and cold air
  • Adjust fan speeds using measured demand
  • Review rack layouts
  • Remove abandoned cabling that blocks airflow

Airflow improvement can reduce cooling energy without changing equipment temperature limits or redundancy.

Containment should be designed around failure behavior

Hot aisle or cold aisle containment reduces mixing between supply and return air.

This can improve:

  • Cooling efficiency
  • Temperature consistency
  • Cooling capacity
  • Rack density
  • Predictability

However, containment changes how the room behaves during:

  • Cooling failure
  • Fan failure
  • Door opening
  • Power interruption
  • Fire events
  • Maintenance
  • Sensor failure

Before deployment, evaluate:

  • Emergency airflow
  • Pressure balance
  • Fire suppression
  • Door control
  • Leak paths
  • Human access
  • Temperature rise rate
  • Sensor placement
  • Backup cooling
  • Failure alarms

Containment should improve normal efficiency while preserving a safe response during abnormal conditions.

The guide to data center cooling systems outlines the major cooling approaches and the operational considerations behind them.

Raise temperature setpoints based on evidence

Increasing supply air temperature can reduce cooling energy.

It may enable:

  • More economizer hours
  • Higher chiller efficiency
  • Lower compressor load
  • Reduced fan energy
  • Better use of ambient conditions

However, setpoint changes should be based on equipment inlet temperature, not only room temperature.

A safe process includes:

  1. Verify sensor accuracy.
  2. Place sensors at representative equipment inlets.
  3. Identify the hottest racks.
  4. Review vendor temperature limits.
  5. Check redundancy and cooling capacity.
  6. Increase setpoints in small steps.
  7. Observe equipment temperatures and alarms.
  8. Test during high load and hot weather.
  9. Document rollback criteria.
  10. Continue trend monitoring.

A room level sensor may show an acceptable value while upper rack equipment receives much warmer air.

The goal is to avoid overcooling while maintaining acceptable inlet conditions for every critical device.

Use variable speed equipment where possible

Fans and pumps often consume less energy when speed is reduced.

Many cooling systems can adjust output according to actual demand through:

  • Variable frequency drives
  • Electronically commutated fans
  • Variable speed pumps
  • Dynamic pressure control
  • Temperature based control
  • Load based control

This can reduce energy compared with running equipment continuously at full speed.

Control changes must be tested carefully.

Risks include:

  • Slow response to rapid load changes
  • Incorrect sensor input
  • Poorly tuned control loops
  • Minimum flow limits
  • Uneven pressure
  • Sensor failure
  • Communication failure

Safe optimization requires monitoring, alarms, minimum operating limits, and fallback behavior.

Match cooling output to the real heat load

Overcooling often occurs when cooling systems operate according to static assumptions.

A better approach uses real time information such as:

  • Rack power
  • Inlet temperature
  • Return temperature
  • Air pressure
  • Humidity
  • Cooling unit load
  • IT workload
  • Weather
  • Chilled water temperature
  • Flow rate

This data can help operators reduce unnecessary cooling while maintaining thermal margin.

The most effective control point depends on the facility design.

For example:

  • Room based cooling may use zone temperature and pressure
  • In row cooling may respond to local rack load
  • Rear door heat exchangers may respond to water and exhaust temperature
  • Liquid cooling may respond to flow, pressure, and coolant temperature

Automation should remain transparent. Operators need to understand what changed, why it changed, and how to return to a safe state.

Expand economizer use when conditions allow

Air side and water side economizers use favorable outdoor conditions to reduce mechanical cooling.

Potential benefits include:

  • Lower compressor energy
  • Reduced chiller operation
  • Lower cooling cost
  • Improved seasonal PUE

Practical constraints include:

  • Climate
  • Air quality
  • Humidity
  • Contamination
  • Water availability
  • Control complexity
  • Maintenance
  • Local regulations

Economizer performance should be evaluated across the full year.

A design that works well in a cool, dry climate may be unsuitable in a hot, humid, polluted, or water constrained location.

Reliability controls should include:

  • Environmental monitoring
  • Filtration
  • Humidity control
  • Automatic mode switching
  • Alarm thresholds
  • Backup mechanical cooling
  • Maintenance procedures

Reduce power conversion losses

Energy is lost as power moves through the facility.

Losses may occur in:

  • Transformers
  • UPS systems
  • Batteries
  • Power distribution units
  • Cables
  • Power supplies
  • Voltage conversion stages

Efficiency improvements can include:

  • Operate UPS systems within efficient load ranges
  • Consolidate lightly loaded UPS modules
  • Use efficient transformer designs
  • Reduce unnecessary conversion stages
  • Improve voltage matching
  • Balance phases
  • Maintain clean electrical connections
  • Replace inefficient legacy equipment
  • Review distribution topology

Redundancy requirements must remain intact.

For example, consolidating UPS modules may improve efficiency, but the final configuration must still support the agreed failure and maintenance scenarios.

Efficiency cannot be assessed only during normal operation. It must also be assessed during component failure, bypass, testing, and maintenance.

Eliminate ghost load and abandoned equipment

Unused or underused equipment still consumes power and creates heat.

Examples include:

  • Powered but idle servers
  • Abandoned network equipment
  • Old storage systems
  • Duplicate appliances
  • Unused development systems
  • Spare devices left online
  • Legacy equipment after migration
  • Empty chassis with active components

Removing this equipment can reduce:

  • IT energy
  • Cooling load
  • Rack usage
  • Port usage
  • Maintenance effort
  • Monitoring noise

However, there is an important PUE effect.

Removing IT load can sometimes make PUE appear worse because the denominator becomes smaller while fixed facility loads remain.

This does not mean the action was harmful.

Total energy and operating cost may still decline significantly.

That is why PUE should never be used alone.

Improve IT utilization as well as facility efficiency

A facility can have an excellent PUE while supporting poorly utilized servers.

IT efficiency improvements may include:

  • Consolidate workloads
  • Retire unused systems
  • Improve virtualization
  • Use power management
  • Schedule noncritical workloads
  • Match hardware to workload
  • Refresh inefficient equipment
  • Improve storage efficiency
  • Reduce duplicate environments
  • Use autoscaling where appropriate

These measures reduce energy per useful unit of computing.

They may not always improve PUE because they reduce IT energy, but they can improve the overall environmental and financial outcome.

A balanced program should track both facility efficiency and IT productivity.

Use device level power data for placement decisions

Rack placement affects cooling and power efficiency.

If high density equipment is concentrated without planning, the facility may need excessive airflow or lower temperature setpoints to protect one hotspot.

Better placement considers:

  • Actual device power
  • Rack power capacity
  • Cooling capacity
  • Airflow
  • Redundant feeds
  • Weight
  • Network requirements
  • Future growth
  • Maintenance access

Spreading load evenly may help some facilities. Concentrating high density equipment into purpose built zones may help others.

The correct strategy depends on cooling architecture and distribution design.

A data center power calculator can support initial rack estimates, while real device telemetry should guide ongoing decisions.

Optimize humidity control without creating instability

Legacy humidity control strategies can consume significant energy.

Facilities may humidify and dehumidify at the same time if systems are poorly coordinated.

Improvement opportunities include:

  • Review acceptable humidity ranges
  • Calibrate sensors
  • Coordinate cooling units
  • Reduce overlapping control
  • Use dew point based control where appropriate
  • Improve control deadbands
  • Review seasonal behavior
  • Eliminate unnecessary humidification

Changes should protect equipment from:

  • Condensation
  • Static risk
  • Corrosion
  • Rapid environmental change

Humidity control should be optimized carefully and verified across the facility, not adjusted based on one sensor.

Maintain equipment to preserve efficiency

Dirty, worn, or poorly calibrated equipment consumes more energy.

Maintenance actions that support efficiency include:

  • Clean filters
  • Clean coils
  • Inspect heat exchangers
  • Calibrate sensors
  • Repair leaking valves
  • Maintain pumps
  • Check fan performance
  • Verify refrigerant levels
  • Inspect dampers
  • Maintain cooling towers
  • Test controls
  • Review alarm history

Maintenance also protects reliability.

Efficiency projects often focus on new technology, but restoring existing systems to proper condition may deliver faster and safer results.

Avoid disabling redundancy for a better number

Some efficiency measures can reduce redundancy.

Examples include:

  • Turning off redundant cooling units
  • Reducing active UPS modules
  • Closing backup airflow paths
  • Raising setpoints without thermal margin
  • Reducing pump capacity
  • Operating near circuit limits
  • Reducing spare capacity

These measures may lower PUE during normal operation, but they can increase risk during:

  • Equipment failure
  • Maintenance
  • Utility disturbance
  • Sudden workload increase
  • Extreme weather
  • Control failure

A safe optimization process should define the minimum required resilience.

This may include:

  • N plus one cooling
  • Dual power paths
  • UPS reserve
  • Generator reserve
  • Temperature recovery time
  • Spare capacity
  • Maintenance tolerance
  • Failure response

The efficiency target must operate inside these limits.

Test changes under realistic conditions

A change that works during low load may fail during peak demand.

Testing should include:

  • Peak IT load
  • Hot weather
  • Low facility load
  • One cooling unit unavailable
  • One power path unavailable
  • Maintenance mode
  • Rapid workload change
  • Sensor failure
  • Communication failure
  • Emergency response

Not every test requires a live failure. Modeling, staged tests, and controlled maintenance windows can provide useful evidence.

The important point is that optimization should be validated beyond ideal conditions.

Use change control for energy optimization

Efficiency changes should follow the same discipline as other production changes.

A strong process includes:

  • Baseline measurement
  • Defined objective
  • Risk review
  • Affected equipment list
  • Success criteria
  • Alarm thresholds
  • Rollback plan
  • Change window
  • Monitoring period
  • Final review

For example, a temperature setpoint increase should specify:

  • Current temperature
  • Proposed temperature
  • Expected energy effect
  • Maximum equipment inlet temperature
  • Alarm threshold
  • Test duration
  • Rollback condition
  • Responsible owner

This turns efficiency improvement into a controlled operational program.

Track several metrics together

PUE is more useful when combined with other measures.

Recommended measures include:

  • Total facility energy
  • IT energy
  • Cooling energy
  • Electricity cost
  • Peak demand
  • Rack power
  • Server utilization
  • Temperature compliance
  • Number of thermal alarms
  • Cooling reserve
  • UPS efficiency
  • Water use
  • Carbon intensity
  • Workload output
  • Service availability

A change should be considered successful when it reduces waste without causing:

  • More alarms
  • Higher equipment temperature risk
  • Reduced redundancy
  • More downtime
  • More maintenance
  • Unstable controls
  • Lower workload performance

Prioritize improvements by risk and payback

Not every facility needs the same sequence.

A practical order often begins with:

  1. Correct measurement.
  2. Remove airflow problems.
  3. Calibrate sensors.
  4. Repair inefficient equipment.
  5. Optimize setpoints gradually.
  6. Tune fan and pump control.
  7. Improve containment.
  8. Remove unused IT equipment.
  9. Improve placement and utilization.
  10. Evaluate major cooling upgrades.

This order starts with relatively low risk measures before moving toward larger capital projects.

Each action should be evaluated for:

  • Energy reduction
  • Cost
  • Implementation effort
  • Operational risk
  • Payback period
  • Reliability effect
  • Maintenance effect

PUE improvement is a continuous operating practice

There is no permanent final PUE.

The number changes as:

  • Workloads grow
  • Equipment is refreshed
  • Rack density changes
  • Weather changes
  • Cooling systems age
  • Sensors drift
  • Facilities expand
  • Operating practices change

Teams should review efficiency regularly and connect it to capacity, maintenance, and lifecycle planning.

The most effective programs combine:

  • Accurate measurement
  • Facility telemetry
  • IT power data
  • Environmental monitoring
  • Capacity data
  • Change management
  • Maintenance
  • Business priorities

Reliability is the boundary condition

PUE is valuable because it exposes facility overhead.

It helps teams find waste in cooling, power conversion, airflow, and supporting systems.

But a data center exists to provide reliable computing.

An efficiency change that increases outage risk, thermal instability, or maintenance exposure can create costs far greater than the energy saved.

The right target is therefore not the lowest theoretical PUE.

It is the lowest sustainable PUE that the facility can achieve while maintaining its agreed resilience, environmental limits, and service obligations.

That balance produces real efficiency rather than a better number alone.

Originally published on the Sensaka blog.

Top comments (0)