Summary: A UPS earns its rating by surviving a load step. The standard behind that rating, IEC 62040-3, treats output stability as a one-time test. An AI training run asks the same unit to perform that step several times a second, for hours. Nothing about that is non-compliant, and nothing about compliance tells you how your own modules are coping. The number worth knowing is not the rating on the nameplate, it is the duty cycle your power train is actually running.

AI loads behave differently

Traditional compute is, from an electrical point of view, relatively stable. Peaks arrive with some warning. This means UPS systems are not having to handle sudden large spikes in load.Large-scale AI training does not behave the same way. According to Uptime Institute, the gap between the low and high points of power draw during a training run can double almost instantaneously, within milliseconds. Uptime puts worst-case swings at 150%, with overshoots more than 10% above system specification.

At rack level, they give a specific example. An Nvidia GB200 NVL72 can climb from 60 to 70 kW up to more than 150 kW. That upper figure is roughly 20% above the rack’s own 132 kW maximum specification. At facility scale, loads move from 45% to 50% of design capacity to 80% to 85% within a second.The reason those swings do not soften as they travel upstream is synchronisation. A bulk training job has every GPU computing and pausing together, so the waveform is closer to a square wave than a ramp, and it repeats. There can be spikes every second or two during batch loading. Thousands of nodes doing that in step do not average each other out. Instead, they compound as they travel upstream.

The standard tests one step. The workload delivers millions.

Vertiv published testing of its large UPS frames against simulated AI load patterns. The conditions are worth reading closely:

ScenarioPatternFrequencyLoad stepsDuration
1100 ms on, 30 ms off8.3 Hz0 to 50%, up to 0 to 100%4 hours
2500 ms on, 500 ms off1 Hz30 to 50%, up to 30 to 100%4 hours
3100 ms on, 900 ms off1 Hz30 to 50%, up to 30 to 100%4 hours

The results appear good. Vertiv reports the tested systems handled steps from 0 to 100% of rated capacity without drawing on the battery, as long as utility supply stayed within nominal parameters. Do note, however, that this is the manufacturer testing its own products.Vertiv notes that IEC 62040-3, the UPS performance standard, treats output stability as a one-time test. A unit can be entirely compliant on the strength of a single load step, then spend a four hour training run performing that step several times a second. These tests both answer different questions. One asks whether the unit can do it. The other asks what happens on the ten-thousandth repetition.

This is a mismatch between how equipment is qualified and how the new class of AI workloads uses it, and mismatches like that are normally discovered in operation rather than in a specification.

The loads are concentrated, not distributed

AI racks cluster. They are deployed in pods, share networking topology, and are frequently placed together for cooling reasons. That means the load dynamics described above do not distribute themselves evenly across a facility’s power train. They land on whichever modules happen to feed the GPU rows.

So two facilities with the same total IT load and the same UPS capacity can have completely different risk profiles, depending on whether the AI racks sit behind one module or four. That is a design decision, not a capacity question, and a capacity report at the facilities level doesn’t show it. Uptime’s own guidance points the same direction when it recommends N+2 redundancy for these environments specifically to avoid repeated overloads landing on the units still in service.

What the rack-level fixes do, and what they leave behind

The industry is responding to this, mostly upstream of the UPS.NVIDIA’s GB300 NVL72 puts capacitors in the power shelves. They charge during low demand and discharge into the peaks, which NVIDIA says reduces grid-facing peak power by 30%. Two further protection mechanisms are also implemented. A power cap that gradually ramps GPU draw up at workload start so the ramp stays inside grid tolerance, and a controlled “GPU burn” at the end that briefly holds draw high so the ramp-down tapers rather than collapsing. Ramp rates and idle thresholds are configurable through nvidia-smi or Redfish.

This genuinely helps spread the loads and reduce sudden load surges. If your GB300 deployment is smoothing the curve at its interconnection, that is good news for your utility relationship and reduces impact on your UPS systems.

So what should I be measuring?

The honest answer to “is my UPS coping” is not available from its spec sheet, a compliance certificate or a monthly average. All three describe a design. What an operator needs is data on its behaviour.How often, and how deep. Not the peak, which everyone records, but the frequency and depth of peaks over a training run. A module that touches 85% load once a week is in a different situation from one touching 85% every ninety seconds. The important factor here is that the sampling resolution decides what you can know. You can read more about this in the polling interval piece.

Which racks sit behind which module. This is the concentration problem, and it is answerable only with a model of the power train itself, from mainline through UPS to PDU and outlet. Without that map, an operator can see a module under stress and still not know which workloads are causing it.Whether the pattern is changing. Duty cycle is a function of the job mix. A facility that adds a training cluster, or changes batch size, changes its electrical behaviour without changing anything the capacity plan tracks.

AKCP monitors UPS systems as virtual sensors and models the full power train in Quicklime, from mainline power through the UPS and PDU layer down to individual outlets, which is what makes the concentration question answerable. Two honest boundaries on that. Our sampling floor is one second, so we resolve the second-scale envelope Uptime describes and not the individual 100 ms transient inside it. And we are a monitoring platform, not a power quality analyser, so harmonics and waveform analysis belong to different instruments.

That is a narrower claim than “we solve AI power dynamics”, and it is the one that survives contact with an electrical engineer. The envelope and the topology are the two things an operator can act on, and most facilities currently have neither.

UPS Monitoring Solutions

Utilizing software such as AKCP’s Quicklime DCIM you can monitor your complete UPS stack. Integration via SNMP to your UPS system and overlaying this data onto your power train, mapped out in Quicklime, you can see live data, as well as historical data. Graph your workloads and spot spikes, this helps you manage your loads, protect your backup power systems and help with load balancing of AI workloads across your power train.

Frequently asked questions

Does AI actually damage a UPS, or is this a planning concern?

Neither the Uptime nor the Vertiv material claims damage in normal operation. Vertiv’s tested frames handled full-range steps without battery involvement. The concern is that qualification is based on a single event while operation involves continuous repetition, and that the consequences of repetition are not established by the compliance test.

Is IEC 62040-3 inadequate?

The standard does what it was designed to do. It predates workloads that cycle a UPS at these frequencies for hours. A one-time output stability test is a reasonable qualification for loads that step occasionally, which described almost everything until recently.

Will the NVIDIA GB300 power shelf solve this?

It reduces grid-facing peak power by 30% according to NVIDIA, which addresses the utility-facing half of the problem. The GPUs still draw the same pattern, so equipment between the rack and the meter still sees the cycling.

What sampling rate is needed to see this?

To resolve the second-scale envelope, seconds. To resolve an individual 100 ms transient or an 8.3 Hz cycle, considerably faster than most infrastructure monitoring platforms operate, including ours. Be sceptical of any monitoring vendor claiming to capture millisecond electrical events as a routine polled metric.

Where should an operator start?

With the power train map. Knowing which UPS modules carry the GPU racks changes how every subsequent number is read, and it is worth establishing before the next capacity conversation rather than during it.

If you want to see what your own power train looks like once it is mapped, and how often your UPS modules are actually being driven, book fifteen minutes.

Sources