Key Takeaways
- DSX MaxLPS enables up to 40% more Vera Rubin GPU capacity within the same power envelope.
- Dynamic Power Software redistributes unused headroom from low-demand racks to active compute in real time.
- 45°C warm-water liquid cooling reduces chiller demand and improves PUE, reclaiming power for compute.
Table of Contents
The Megawatt Math That Just Rewrote AI Factory Design
AI factories face a different bottleneck today. The limiting resource is not physical rack space; it is the volume of useful inference output each megawatt can sustain.
A technical deep dive published by NVIDIA Developer Blog describes a site-level architecture called DSX MaxLPS that could raise Vera Rubin NVL72 GPU capacity by up to 40 percent inside a fixed power envelope.
The implication matters most for inference. Application-level performance per watt has become the primary efficiency metric because token throughput, not raw accelerator count, determines revenue.
Power distribution, cooling, networking, and facility overhead all take their share before a watt reaches compute. The opening challenge is therefore simple: reclaim every watt that does not produce useful AI output.
Inside the DSX MaxLPS Stack: Power Steering, Workload Profiles, and Thermal Headroom
As detailed in the NVIDIA Developer Blog, DSX MaxLPS stands for Maximum Land Power Shell. The name frames an AI factory as a system bounded by land, utility power, and physical shell.
The architecture targets three layers simultaneously.
- Dynamic power allocation — continuously redirects unused headroom from low-demand racks and GPUs to active compute.
- Advanced performance per watt — software profiles align GPU configuration with workload phase to lift output at a fixed power budget.
- 45° C thermal efficiency — warm-water liquid cooling reduces chiller demand and improves PUE.
Traditional data center planning reserves enough power for every rack to draw its peak simultaneously. That approach treats each rack as an isolated island, so spare power cannot flow to a neighbor that needs it.
One representative power-budget view shows only about 60 percent of delivered site power reaching AI compute. Facility overhead, rack losses, and operational inefficiencies consume the rest.
Dynamic Power Software, currently in Developer Preview, replaces that static model. It maps the data center from utility feed down to individual GPUs.
Operators define resource groups, power budgets, and policies. The control loop then continuously compares allocated power with actual consumption.
When a GPU or rack draws below its reservation, DPS makes that headroom available elsewhere in the managed group. The site-level envelope stays unchanged.
An optional event bus called DSX Exchange links DPS to building management systems, electrical monitoring, cooling infrastructure, and compute schedulers. It exposes signals such as seasonal cooling headroom that DPS can act on.
Beyond rack-level steering, MaxLPS includes workload profile power solutions, or WPPS, for inference, training, memory-bound, and compute-bound operating modes. The Application Performance and Power Manager applies these validated profiles to participating GPUs.
NVIDIA Dynamo can further optimize inter-rack power behavior for inference services. The principle is to align GPU configuration, application topology, and serving behavior with fleet-wide output per watt.
Representative inference workload evaluations show why the effort matters. On GB200 NVL72, provisioned rack power dropped from 125 kW to 90 kW, enabling 39 percent more racks in the same envelope.
Vera Rubin NVL72 systems tested with DeepSeek-R1 saw provisioned rack power fall from 136 kW to 101 kW. That supports 35 percent more racks while preserving throughput.
Performance per watt improved roughly 1.5x on GB200 NVL72 and between 1.3x and 1.4x on Vera Rubin NVL72. Combined with data center power planning, the projection reaches 40 percent more Rubin GPU capacity.
In a 100 MW AI factory, that transforms into roughly 40,000 Vera Rubin GPUs. The resulting compute envelope includes 2 zettaflops of NVFP4 inference, 1.4 zettaflops of NVFP4 training, 11 petabytes of HBM4 capacity, and 800 petabytes per second of memory bandwidth.
Thermal design is the third pillar. Vera Rubin NVL72 racks are specified to accept liquid cooling at a 45° C inlet temperature.
Warmer inlet temperatures let facilities lean on free cooling more often. That reduces mechanical chilling and annualized PUE, turning cooled power back into compute.
Many sites run cooling loops colder than necessary because reactive PID control cannot anticipate rapid thermal swings. Agentic control work highlighted by Phaidra uses power telemetry and learned policies to predict those shifts.
Finally, MaxLPS defines a lifecycle capacity target rather than a day-one requirement. Facilities size power, cooling, and network capacity for the full GPU position count upfront, then populate racks as workloads shift toward inference.
The Market Signal: Stranded Capacity Is Now a Boardroom Metric
The DSX MaxLPS announcement changes how AI infrastructure operators should evaluate capital deployment. Efficiency is no longer a facilities afterthought; it is a revenue lever.
A 100 MW site that supports 40 percent more Vera Rubin GPUs without new grid capacity has a materially different payback profile. The gain comes from software and thermal design, not from negotiating additional utility power.
For inference-heavy providers, shifting unused rack headroom to active decoding directly lowers cost per million tokens. That is the difference between a profitable token service and a marginal one.
The cooling strategy may also alter site selection. Facilities designed for 45° C inlet operation can use free cooling in more climate zones, reducing reliance on energy-intensive chillers or evaporative systems.
Still, the software stack remains in Developer Preview. Dynamic Power Software and DSX Exchange have not reached general availability, so production operators should treat the 40 percent figure as a planning target rather than a guaranteed deployment outcome.
The published validations use representative inference workloads on NVIDIA systems. They are not independent third-party benchmarks, and results will vary by model mix, cooling infrastructure, and operational maturity.
That caveat does not diminish the strategic signal. It clarifies that the next wave of AI factory competition will be won by operators who pair hardware deployment with power-aware software control.
The New Operating Model for Power-Constrained AI Infrastructure
The power-constrained AI factory now has a clear operating model: treat every megawatt as a software-steerable, workload-aware asset instead of a static rack boundary. For teams translating AI infrastructure shifts into search-visible technical content, programmatic SEO and AI automation is how Andres SEO Expert approaches it — talk to the team here.
Frequently Asked Questions
What is DSX MaxLPS in NVIDIA AI factory design?
DSX MaxLPS, which stands for Maximum Land Power Shell, is a site-level architecture from NVIDIA designed to maximize AI factory performance per watt. It combines dynamic power allocation, advanced performance-per-watt software profiles, and 45°C warm-water liquid cooling to potentially increase Vera Rubin NVL72 GPU capacity by up to 40 percent within a fixed power envelope.
How does Dynamic Power Software (DPS) improve power efficiency in AI data centers?
Dynamic Power Software (DPS), currently in Developer Preview, replaces the traditional static power reservation model. It maps the data center from utility feed down to individual GPUs, continuously comparing allocated power with actual consumption. When a GPU or rack draws below its reservation, DPS makes that unused headroom available to other racks or GPUs in the same managed group, without changing the site-level power envelope.
What are workload profile power solutions (WPPS) in the DSX MaxLPS stack?
Workload profile power solutions (WPPS) are validated software profiles for inference, training, memory-bound, and compute-bound operating modes. The Application Performance and Power Manager applies these profiles to participating GPUs to align GPU configuration with workload phases, boosting output while staying within the same power budget.
How much additional GPU capacity can DSX MaxLPS deliver in a 100 MW AI factory?
In a 100 MW AI factory, DSX MaxLPS could support roughly 40,000 Vera Rubin GPUs, representing a 40 percent increase in GPU capacity. This translates into 2 zettaflops of NVFP4 inference, 1.4 zettaflops of NVFP4 training, 11 petabytes of HBM4 capacity, and 800 petabytes per second of memory bandwidth.
Why is 45°C liquid cooling important for AI factories?
Vera Rubin NVL72 racks accept liquid cooling at a 45°C inlet temperature, which is warmer than typical cooling loops. Warmer inlet temperatures allow facilities to rely more on free cooling, reducing mechanical chilling and annualized PUE. This effectively turns cooled power back into compute capacity and improves overall energy efficiency.
What is the market significance of DSX MaxLPS for AI infrastructure operators?
DSX MaxLPS signals that efficiency is now a revenue lever, not just a facilities concern. A 100 MW site supporting 40 percent more GPUs without new grid capacity has a different payback profile. For inference-heavy providers, shifting unused rack headroom to active decoding lowers cost per million tokens, which can be the difference between a profitable token service and a marginal one.
What caveats should operators consider before adopting DSX MaxLPS?
DSX MaxLPS components like Dynamic Power Software and DSX Exchange are still in Developer Preview and have not reached general availability. The published validations use representative NVIDIA workloads, not independent third-party benchmarks. The 40 percent capacity figure is a planning target rather than a guaranteed outcome, and actual results will vary based on model mix, cooling infrastructure, and operational maturity.
