Key Takeaways
- NVIDIA DSX MaxLPS ran 192 GPUs inside a 264.4 kW envelope that previously held 140, delivering 49.2% more aggregate inference throughput with no added provisioned power.
- Throughput per provisioned watt rose from 4.10 to 6.12 tokens per second per watt, while power-budget utilization climbed from 62.9% to 75.2%.
- Median and P75 latency stayed within 5% of baseline, but P99 time to first token rose 17% — making tail-latency limits as critical as throughput targets in production acceptance criteria.
Table of Contents
A 264.4 kW Envelope Just Unlocked 37% More GPUs
On September 27, 2026, a renewable-powered data center in Keflavík, Iceland, delivered a 49.2% increase in aggregate AI inference throughput without adding a single megawatt of provisioned power.
The result came from running 192 GPUs inside the same 264.4 kW electrical envelope that previously supported 140.
DSX MaxLPS, the software behind the test, is designed to enable up to 40% more GPUs within the same approved power budget by dynamically reallocating power across participating resources.
NVIDIA’s joint evaluation with Nscale, as described on the NVIDIA Developer Blog, tested Kimi K2.5 workloads on GB300 NVL72 systems and demonstrated how that policy-governed allocation can reclaim capacity that static provisioning leaves stranded.
That shift is now moving from a software feature into a core AI factory procurement metric.
Inside the Policy-Governed Power Loop
Static power planning typically reserves enough headroom for every GPU to peak at the same moment.
Real AI workloads rarely behave that way.
Training runs cycle through compute, communication, synchronization, and checkpointing; inference shifts among prefill, decode, memory-bound, and idle states.
DSX MaxLPS operates as a coordinated allocation layer, not a site power increase.
It maps topology into managed resource groups, gathers telemetry at the GPU, node, rack, and group level, and then adjusts GPU power limits according to operator-defined policy.
When some resources draw below allocation, the control loop reallocates that unused headroom to other nodes inside the same approved boundary.
The control process has five technical elements:
- Topology and resource groups: Operators map participating infrastructure and organize nodes into a managed group with an aggregate power budget.
- Telemetry: Collection spans GPU, node, rack, and group power data at intervals sufficient to detect headroom and emerging power events.
- Policy: Rules define node limits, group limits, allocation priorities, reserve requirements, and responses to maintenance or emergency events.
- Allocation and control: The software adjusts participating GPU power limits when some resources draw less than their allocation.
- Validation and enforcement: Measured power is compared against the approved budget and corrected when consumption approaches a limit.
The evaluation used Blackwell Ultra GPUs, Kimi K2.5 in FP4, NVIDIA Dynamo, TensorRT LLM, an 8K input sequence length, and a 1K output sequence length.
The workload mix combined high-throughput and low-latency inference instances to create distinct power and service profiles within the managed group.
The static baseline used 35 four-GPU nodes, while the DSX MaxLPS configuration used 48 four-GPU nodes.
Total measured power rose from 166.2 kW to 198.9 kW, and power-budget utilization moved from 62.9% to 75.2%.
That higher utilization is the entire point: the site already had usable headroom, but static reservations prevented it from being allocated.
Why Tail Latency Is the Real Tax on Power Sharing
The configuration added a third 52-GPU high-throughput instance while keeping one 36-GPU low-latency instance unchanged.
Aggregate throughput rose from 1,084,503 to 1,618,443 tokens per second, and throughput per provisioned watt climbed from 4.10 to 6.12 tokens per second per watt.
High-throughput output per instance stayed essentially flat at 59,153 versus 59,220 tokens per second, while low-latency output remained at 2,265 tokens per second.
Those per-instance results confirm that the larger managed fleet increased total work without materially degrading existing service throughput.
Median and P75 latency stayed within 5% of baseline.
P99 time to first token, however, rose 17% from the 15.7-second baseline.
That divergence is the central operational warning: stable average latency can hide meaningful degradation for the slowest 1% of requests.
Production acceptance criteria therefore need to define tail-latency limits as explicitly as throughput targets.
Available headroom also depends on workload mix.
Complementary power profiles create more opportunity than workloads that peak simultaneously.
Operators should test representative production workloads against the aggregate power limit before adding capacity.
The Next AI Factory Bottleneck Is the Building Itself
Data Center Frontier has documented the broader DSX push: NVIDIA launched DSX as an AI factory platform at GTC Taipei on May 31, with power optimization tied directly to 45-degree Celsius liquid cooling.
NVIDIA estimates that the combined MaxLPS approach could support up to 40% additional Rubin GPU capacity inside a hypothetical 100 MW AI factory.
Trane and Eaton have announced a DSX-aligned power-and-cooling reference design projecting up to 15% energy-efficiency gains, up to 30% lower installation costs, and up to 80% lower copper requirements compared with conventional low-voltage approaches.
Those supplier moves signal that dynamic power allocation is becoming an infrastructure procurement variable, not just a runtime software feature.
An open-access review in Advances in Applied Energy reinforces the stakes: AI training workloads can swing from idle to peak within milliseconds, while rack-level demand can exceed 100 kW and facility consumption may scale to hundreds of megawatts.
The same review notes that GB300 NVL72 designs already include chip-level power smoothing and firmware-controlled ramp-up, with NVML exposing enforced GPU power limits for coordination with rack and facility envelopes.
Industry analysis tracks rack density climbing from 8.5 kW per rack in 2020 to 25 kW in 2022, with expectations of 1 MW per rack by 2028.
Conventional AC distribution is reaching physical limits beyond roughly 170 kW per rack, which is pushing early work on 800 VDC architectures and sidecar power shelves capable of delivering up to about 400 kW per rack.
AI rack power demand can swing from 30% to 100% utilization within milliseconds, far sharper than the peaks seen in non-AI workloads.
NVIDIA has also made strategic investments in Lancium and Cloverleaf Infrastructure, two partners applying DSX earlier in data center development.
Lancium’s pipeline includes more than 15 GW of powered-land capacity, with 4 GW already under lease.
DSX Flex has already been demonstrated in a commercial multi-megawatt pilot with Emerald AI and Silicon Valley Power, where workloads adjust to utility signals while protecting critical AI jobs.
For future Vera Rubin NVL72 factories, NVIDIA’s MaxLPS also pairs dynamic power management with infrastructure designed for 45°C liquid-cooling inlet operation.
Any Vera Rubin capacity projection remains separate from the measured GB300 NVL72 evaluation.
From Stranded Power to Managed Capacity
The Iceland evaluation proves that power-constrained AI factories can unlock meaningful GPU capacity without waiting on additional grid supply.
The real engineering discipline is not the software alone; it is defining boundaries, measuring representative behavior, and proving compliance under adverse conditions before scaling.
For teams translating AI infrastructure shifts into search-visible authority, programmatic SEO AI automation is how Andres SEO Expert approaches it — contact us here.
Frequently Asked Questions
What is NVIDIA DSX MaxLPS and how does it work?
NVIDIA DSX MaxLPS is a policy-governed power allocation layer that maps topology into managed resource groups, collects telemetry at GPU, node, rack, and group levels, and adjusts GPU power limits according to operator policy. When some resources draw below allocation, it reallocates unused headroom to other nodes inside the same approved power boundary. NVIDIA says it can enable up to 40% more GPUs within the same power budget.
How did the Iceland data center get more GPUs without adding power?
The Keflavik evaluation ran 192 GPUs inside the same 264.4 kW electrical envelope that previously supported 140 GPUs. It delivered a 49.2% increase in aggregate AI inference throughput without adding a single megawatt of provisioned power. The static baseline used 35 four-GPU nodes, while the DSX MaxLPS configuration used 48 four-GPU nodes.
What were the measured power and throughput gains in the DSX MaxLPS test?
Total measured power rose from 166.2 kW to 198.9 kW, and power-budget utilization moved from 62.9% to 75.2%. Aggregate throughput rose from 1,084,503 to 1,618,443 tokens per second, and throughput per provisioned watt climbed from 4.10 to 6.12 tokens per second per watt. The test used GB300 NVL72 systems, Kimi K2.5 in FP4, NVIDIA Dynamo, TensorRT LLM, an 8K input sequence length, and a 1K output sequence length.
What is the main operational risk of dynamic power sharing?
The main risk is tail latency. In the evaluation, median and P75 latency stayed within 5% of baseline, but P99 time to first token rose 17% from the 15.7-second baseline. That means stable average latency can hide meaningful degradation for the slowest 1% of requests. Production acceptance criteria should define tail-latency limits as explicitly as throughput targets.
Does DSX MaxLPS reduce per-instance throughput or degrade existing service?
Per-instance results showed that the larger managed fleet increased total work without materially degrading existing service throughput. High-throughput output per instance stayed essentially flat at 59,153 versus 59,220 tokens per second, while low-latency output remained at 2,265 tokens per second. The configuration added a third 52-GPU high-throughput instance while keeping one 36-GPU low-latency instance unchanged.
Why is the building itself becoming the next AI factory bottleneck?
AI rack power demand can swing from 30% to 100% utilization within milliseconds, while rack density has climbed from 8.5 kW per rack in 2020 to 25 kW in 2022, with expectations of 1 MW per rack by 2028. Conventional AC distribution is reaching physical limits beyond roughly 170 kW per rack, which is pushing early work on 800 VDC architectures and sidecar power shelves capable of delivering up to about 400 kW per rack. NVIDIA is also tying DSX power optimization to 45-degree Celsius liquid cooling.
Is DSX MaxLPS only a software feature or does it affect infrastructure procurement?
DSX MaxLPS is a runtime allocation layer, but it is becoming an infrastructure procurement variable. Trane and Eaton have announced a DSX-aligned power-and-cooling reference design projecting up to 15% energy-efficiency gains, up to 30% lower installation costs, and up to 80% lower copper requirements. DSX Flex has also been demonstrated in a commercial multi-megawatt pilot with Emerald AI and Silicon Valley Power, where workloads adjust to utility signals while protecting critical AI jobs.
