The power shortage that supply can’t fix (unless it’s already plugged in)
Try VOLT Intelligence
Get StartedTry VOLT Cloud
Deploy GPUTable of Contents
- What the 38GW shortfall means for the AI industry
- Why every GPU power fix takes years to sort out
- New nuclear deals
- New gas generation
- New substation and transmission build
- Offshore or international campuses
- The GPU compute that's already energized
- Colocation facilities
- Research labs and universities
- Former cryptocurrency mining operations
- Enterprise and cloud overflow
- Time-to-capacity: New campus vs. existing network
- GPU pricing on VOLT vs. hyperscaler on-demand
- Real-world cost comparison: 6-month training run
- Why VOLT for your already-powered GPU needs?

The current AI infrastructure conversation has settled on one big number: 38 gigawatts. That's Morgan Stanley's estimated US data-center power shortfall through 2028. What’s creating this shortfall? The GPU compute needed for inference workloads, and what the existing grid can actually deliver to new AI data facilities on any realistic timeline.
Here are some quick numbers. Grid interconnection queues are presently running 5+ years out in most major markets. Add another 2–3 years for transformer lead times, and new gas turbine agreements that take approximately 18–36 months to work their way through permitting and construction. Add to this the fact that Microsoft, Google, and Amazon signed nuclear deals for power that won't realistically go online until the early 2030s. All of that math adds up, and won’t do any AI startups and LLM research teams any good in the short or long-term.
Every proposed fix to this compute shortage relies on one faulty assumption: build more power capacity. The reality is that a material portion of the compute the industry needs is already energized, cooled, and connected. The problem is that it's pointed at the wrong workload or sitting partially idle.Below, we’ll get into how this power shortfall happened, what it means for your AI project, and how you can tap into idle capacity right now without waiting for power infrastructure builds.
What the 38GW shortfall means for the AI industry
Morgan Stanley's figure represents the mathematical difference between AI expansion plans filed with utilities and the generation and transmission capacity those utilities can actually provision.
Grid interconnection queues in PJM, MISO, and WECC have grown from months to years over the past four years. Lawrence Berkeley National Laboratory's 2024 interconnection study found over 2,600 GW of generation and storage projects waiting in US queues. The majority of these projects will never reach commercial operation. Data centers compete for the same slots as solar farms and battery storage.
Transformer availability compounds this problem. For example, take a 500 MVA power transformer, which is the kind a new hyperscaler campus needs. It currently carries a 2–3 year lead time from major manufacturers. It’s clear that global transformer manufacturing capacity hasn't yet scaled to match demand. Each new campus announcement kicks off a procurement race for the same finite pool of equipment.
There is also a physical damage angle brought on by infrastructure wear and tear. AI workloads—specifically large model training jobs—create power draw profiles that are volatile in ways that legacy data-center infrastructure wasn't designed for. GPU clusters create load spikes that stress UPS batteries, diesel generators, and cooling systems. Facilities designed for more predictable HPC or enterprise workloads are seeing accelerated wear on infrastructure originally specified for 15–20 year lifespans.
If you're a hyperscaler announcing a new AI campus today, you're looking at a 4–7 year path from announcement to meaningful capacity. Site selection, environmental review, utility agreements, interconnection queue, substation construction, transformer procurement, facility construction, and commissioning stack sequentially.
The AI industry’s chip shortage taught everyone that lead times multiply quickly. The power shortage is doing the same thing. This time the underlying constraint, grid infrastructure, doesn't respond to capital the way chip fabs do.
Why every GPU power fix takes years to sort out
If you follow the trade press, solutions fall into a few categories and share the same multi-year expectation.
New nuclear deals
The tech world, mostly in Silicon Valley, has been enamored with nuclear power for several years now. Microsoft's agreement with Constellation, announced in 2024, to restart Three Mile Island's Unit 1 is now the template. The plant is expected to come back online in 2028—a relatively fast turnaround given that this facility already existed. Other new nuclear power plant builds won’t be so lucky on their go-live dates. Agreements for new small modular reactors (SMRs) from vendors like Oklo or X-energy are targeting commercial deployment in the 2028–2032 range, and that’s on the optimistic end of spectrum.
New gas generation
Hyperscalers signing agreements with gas turbine operators face 18–36 month delivery queues for new turbines, plus permitting. Some are leasing mobile generation units to bridge gaps. This approach will work at small scale but it won’t address campus-scale demand.
New substation and transmission build
This is where the interconnection queue problem lives. Building a new substation to serve a 500MW campus requires utility cooperation, ROW acquisition, environmental review, and construction. All told, this typically 3–5 years in the US, and much longer in markets with more constrained permitting.
Offshore or international campuses
Moving workloads to Iceland, Norway, or Finland to access cheap renewable power solves one problem (cost, carbon). However, it also introduces issues like latency, data sovereignty, and regulatory complexity, making its viability for latency-sensitive inference at scale questionable.
There is a common thread running through all of these infrastructural stumbling blocks: every fix treats power as something to generate, route, or build more of. None of them can compress the multi-year timeline that grid infrastructure operates on.
The GPU compute that's already energized
Thankfully, there is a large, distributed inventory of GPU compute that is already energized, cooled, and past its own interconnection process. It's existing capacity that simply needs demand routed to it.
This is exactly what we designed VOLT to do. VOLT's GPU supply aggregates from several source categories, which gives it several built-in advantages, from state-of-the-art silicon to pricing, zero waitlists, and more.
Colocation facilities
This is a jargony way of describing AI data centers that house third-party equipment. Such facilities went through their own interconnection and build-out years ago. That means there’s a good chance that the GPUs they house may be underutilized during off-peak windows, or owned by operators who deployed them speculatively and need revenue. Everyone wins.
Research labs and universities
HPC clusters purchased for specific grant-funded projects often have idle capacity between compute-intensive runs. The power and cooling is paid for whether or not the GPUs are running jobs. When you use GPU clusters from the academic world, you’re also giving universities and their students a nice financial assist in the process, and possibly even supporting LLM research. Again, everyone wins.
Former cryptocurrency mining operations
This is the most distinctive part of VOLT's supply story. Crypto mining facilities were purpose-built for high-density, continuous compute. Indeed, they're optimized for exactly the kind of thermal and power environment that GPU training and inference needs. Many of these facilities went through their interconnection processes from 2019 to 2022, when crypto was booming and the mining economics made large power agreements worthwhile. As mining economics shifted, significant GPU inventory became underutilized in already-energized, already-cooled facilities.
Enterprise and cloud overflow
These are enterprises or organizations with their own GPU clusters that aren't running at full utilization. It makes business sense for them to make that compute available on distributed GPU marketplaces like VOLT.
The argument is not that VOLT is exempt from the power bottleneck. It's more specific than that: VOLT is not asking for new power. It is routing demand to existing energized capacity that is underused right now.
Time-to-capacity: New campus vs. existing network
Let’s look at the two distinct paths to the same goal: getting a thousand H100s running AI workloads. We’ll see what that looks like across a new hyperscaler campus, new colocation deal, and an existing node onboarded to VOLT.
The hyperscaler timeline comes from publicly reported project durations. Look at Amazon's announced Virginia data center expansions, or Google's reported interconnection queue experiences in Texas, as well as Microsoft's public disclosures about its nuclear power agreements. For large-scale new builds, you’re looking at 4–7 years when site selection, transformer builds, utility agreements, and so on are factored into the equation.
Now, granted, VOLT's aggregate power draw doesn't rival a hyperscaler's–which we’re not trying to do, by the way. What we are interested in is the marginal time-to-capacity for the next increment of compute. If an existing mining facility already has 2,000 GPUs connected to power that was provisioned three years ago, the marginal cost of routing a new AI workload to those GPUs doesn’t have to contend with power agreements, interconnection applications, or 4–7 year waits.
GPU pricing on VOLT vs. hyperscaler on-demand
The power issue connects directly to cost, given that the infrastructure provisioning cost is a primary driver of GPU cloud pricing. Here’s a table to visualize the comparison between GPU pricing on VOLT and on-demand pricing you’ll almost certainly find from hyperscalers.
Sure, that’s some mighty fine GPU savings, but there is a caveat here. In the interest of full transparency, we aren’t looking at exactly equivalent product offerings.
Hyperscaler pricing includes SLAs, support tiers, and integrated tooling, amongst some other features that VOLT simply doesn't offer in the same form. But the order-of-magnitude difference in H100 pricing reflects, in part, the capital cost of the infrastructure path. VOLT providers operating in already-amortized facilities don't carry the cost of new grid interconnection, new substations, or new transformer procurement. That cost structure flows through to pricing.
Real-world cost comparison: 6-month training run
A mid-size AI lab running a 70B parameter model fine-tuning job needs approximately 64 H100s for 180 days. This, of course, is a rough estimate for a substantial fine-tuning job, not full pre-training.
On AWS (p4de.24xlarge, 8x H100 equivalent)
8 nodes × $32.77/hr × 24hr × 180 days = $1,139,673
On VOLT at ~$2.09/hr per H100
64 GPUs × $2.09/hr × 24hr × 180 days = $578,995
That works out to approximately $560,000 in savings over six months; and that’s on a job where AWS can actually provision the capacity. It’s absolutely crucial to understand that during periods of high demand or in regions with interconnection constraints, AWS availability for 64 H100s is never guaranteed.
VOLT, by comparison, draws from a geographically distributed pool. This provides a different kind of availability risk profile. You’ll never find your project trapped in a single-vendor queue. Instead, you’ll experience variable hardware quality and uptime SLAs across a distributed set of providers.
Why VOLT for your already-powered GPU needs?
If you’re looking at VOLT to meet your GPU needs that are already powered and not queued, then you already understand the marginal time-to-capacity problem. AI teams like yours simply cannot wait for new data center builds and, even more importantly, cannot absorb hyperscaler pricing.
Remember, our network aggregates GPUs from providers who are already through the hard part of the infrastructure problem. They aren’t sitting in queues for interconnections or transformer and substation builds. VOLT GPU providers already sorted that stuff out.
VOLT’s key contribution here is the software layer: workload matching, hardware verification, payment infrastructure, and the marketplace that routes demand to existing supply.
Again, this doesn’t mean that VOLT is always equivalent to a hyperscaler for every use case. Latency-sensitive applications need to know where in the network their compute actually lives. Compliance-sensitive workloads need attestation of facility jurisdiction and physical security that a distributed network makes harder to guarantee uniformly. For that, hyperscaler compute might make sense for some AI startups and enterprises. VOLT doesn’t claim that kind of uniformity, as node quality and uptime SLAs vary across the network.
What VOLT offers that AWS, Azure, and GCP cannot offer in the near term is access to AI compute that does not depend on new power infrastructure coming online. The 38GW shortfall is a severe constraint on new capacity. For teams running experiments, iterating on models, or running inference at scale where burst availability matters more than single-vendor guarantees, that distinction is practically significant right now, not in 2028 when the next wave of data center builds might come online.
Save up to 70% on GPUs today, without the wait.
Related Questions
What is the AI data center power shortage and when does it resolve?
The AI data center power shortage refers to the gap between the electricity demand created by large-scale GPU compute — training and inference workloads — and the power that grid operators can actually deliver to new facilities. Morgan Stanley estimated a 38GW US shortfall through 2028. The timeline is driven by grid interconnection queues (5+ years in many markets), transformer lead times (2–3 years), and permitting for new generation. Industry analysts don't expect meaningful relief before 2029–2030 at the earliest, and that assumes planned nuclear and gas agreements proceed on schedule. The shortage is very much about the pipeline for getting new power to new data center sites.
How does the grid interconnection queue affect GPU cloud availability?
The interconnection queue determines when a new data center can physically connect to the grid at the capacity it needs. In PJM and MISO, the two largest US grid operators by territory, average queue wait times have grown from under a year to 5+ years since 2020. This means a hyperscaler that breaks ground on a new campus today cannot get full power until the late 2020s at the earliest in constrained markets. That timeline directly limits how fast cloud providers can expand H100 and future GPU capacity in those regions. It’s also one structural reason GPU availability from major cloud providers has been inconsistent despite significant capital investment.
Why is idle GPU capacity relevant to the AI compute bottleneck?
A significant portion of GPU hardware deployed globally is not running AI workloads continuously. Former cryptocurrency mining facilities built during 2019–2022 represent energized, cooled compute infrastructure where the power agreements are already in place and the hardware may be underutilized. Research institutions and colocation tenants similarly have burst capacity available during off-peak windows. Routing AI workloads to this idle inventory only requires software coordination instead of power infrastructure. Decentralized GPU networks like VOLT are built around exactly this routing problem: connecting AI compute demand to existing supply without waiting for new grid capacity.
Does decentralized GPU compute have the same power constraints as hyperscalers?
No — but for a specific reason. Decentralized networks like VOLT don't build data centers; they aggregate capacity from facilities that already exist and are already energized. This means they don't face the same interconnection queue timeline that a hyperscaler building a new campus faces. They do face the constraint that their total available supply is bounded by what existing providers have deployed. They can't unilaterally add 500MW of new capacity the way a hyperscaler with capital and long-term utility contracts can. The advantage is marginal time-to-capacity for the next increment of compute using existing infrastructure. The limitation is that total network scale and reliability guarantees differ from a centralized provider.
What should AI teams consider when evaluating GPU cloud options given power constraints?
The most useful question is "where is my workload in the provisioning queue?" For latency-tolerant training and fine-tuning jobs, distributed networks with existing infrastructure offer faster access to capacity and substantially lower pricing than hyperscaler on-demand. For latency-sensitive inference at production scale, the relevant questions are node location, uptime SLA specifics, and what happens during a node failure. The power bottleneck affects all providers to the extent they depend on new build-outs; providers drawing from existing infrastructure are insulated from that constraint on the supply side while carrying a different risk profile around consistency and SLA enforcement.