The Token Cost Paradox: Why cheaper inference doesn't mean cheaper AI bills
Try VOLT Intelligence
Get StartedTry VOLT Cloud
Deploy GPU
The AI budget conversation is the hot topic this year. So, it’s safe to say that If you've probably heard some version of this line: token costs are falling fast and AI is about to get a lot cheaper.
It's a reasonable proposition. You want to believe it. But according to Gartner's own research, it’s a notion that is mostly wrong for the people actually paying the bills.
Gartner's forecast, published this spring, makes some other striking proclamations. By 2030, they believe that running inference on a large language model with a trillion parameters will cost providers more than 90% less than it did in 2025. But let’s do some math to see if that checks out. If we stretch the window back to the earliest trillion-parameter-class models from 2022, Gartner’s estimates a 100x improvement in cost efficiency. Okay, well, what exactly are the drivers of that efficiency? The usual suspects: better semiconductors, smarter model architectures, higher chip utilization, more inference-specialized silicon, and a growing role for edge devices handling narrower tasks locally instead of routing everything back to a data center.
Gartner, however, is leaving one crucial bit of information out of the equation: these savings won’t be passed on to enterprise customers. It has nothing to do with either pricing strategy or vendor margin-grabbing. Instead, the math gets more paradoxical the closer you look at it. While your AI startup or research team might well experience genuine, measurable token-cost deflation, you might also see your total AI bill climb month-over-month. While no one is to blame for this escalating AI infra bill, there is a solution to it. Below, we’ll explore the bill paradox’s origins, the visibility problem that obscures it, and why cheap tokens aren’t the solution. We’ll show you where you can find savings, and set your startup or research project up for sustained success.
The billing paradox: an origin story
Your CFO may not like to hear it, but this is the crux of the matter: as cost-per-token falls, the total number of tokens required for each task rises even faster. Why? Cheaper tokens unlock capabilities that weren't economical to run before.
Agentic models — that is, the ones making decisions, chaining tool calls, and working through multi-step reasoning — require somewhere between five and thirty times more tokens per task than a standard generative chatbot exchange. Sure, cheaper tokens make agentic workflows viable at scale, but your project will burn through even more tokens and savings, instead of realizing savings through what you’d expect to be a consumption plateau. Remember, agents run more tasks, many of them quite complex, and do so autonomously, leaned to faster token consumption than the per-token-price can realistically fall.
Let’s revisit Gartner’s forecast, as its own framing of this issue is worth considering.
First, we must grasp that falling commodity token prices are not the same thing as advanced reasoning becoming cheap or widely accessible. Gartner’s forecast splits its modeling into two scenarios: 1) a "frontier" scenario built around leading-edge silicon, and, 2) a more conservative "blend" scenario reflecting a realistic mix of hardware most organizations typically run. The blend scenario’s costs are considerably higher than in the frontier one, which has major implications for mid-sized enterprises that realistically won’t be getting frontrunner access to whatever chip Nvidia or a hyperscaler ships next. Instead, their workloads will be running on whatever mix of hardware its cloud provider happens to have available.
The visibility problem beneath the pricing paradox
G. K. Chesterton—a 19th century author who loved a good paradox—once noted: “It isn’t that they can’t see the solution. It is that they can’t see the problem.” Doubly so in AI infrastructure provisioning and pricing.
To wit, the pricing paradox hits especially hard when organizations simply don't have a clear picture of what they're actually spending in the first place. As it turns out, that happens to be a big chunk of them.
Research from VentureBeat's Pulse team found that 44% of enterprises can't rigorously track their AI compute costs at all. We’re not talking about rough estimates here—we’re talking about tracking them with any sort of precision. It’s reasonable to estimate that about half the startups and enterprises making infrastructure and budget decisions are doing so blindly.
IDC's most recent FutureScape research adds some more perspective. They found that organizations with more than a thousand employees and dedicated FinOps teams — in other words, the companies that are supposed to have this all figured out — are still underestimating their AI infrastructure costs by up to 30%. The average enterprise AI budget has grown from roughly $1.2 million a year in 2024 to around $7 million in 2026, while some Fortune 500 companies are now reporting monthly AI inference bills in the tens of millions of dollars.
Granted, some enterprises have the visibility to see this problem, and they’re content with the spend because they don’t want to fall behind competitors. But even for larger companies with big teams and budgets, those numbers simply aren’t sustainable in the long term.
AI utilization rates compounds the problem. Cast AI's 2026 research found that average GPU utilization across surveyed organizations is as low as 5%. Separate research on AI infrastructure at scale found that roughly two-thirds of organizations report peak utilization below 70%, even during their busiest periods. In other words, that's capacity that's been paid for and is sitting idle most of the time. It’s a cost that has nothing to do with the token price and everything to do with how that capacity was provisioned and orchestrated in the first place. A company can be enjoying the full benefit of Gartner's projected 90% token cost decline and still be bleeding money through GPUs that are running at a fraction of their capacity around the clock.
Why cheap tokens don’t simplify procurement
The instinct, understandably, is to shop for AI infra on price-per-token or price-per-GPU-hour, because that's the number that's easiest to compare across providers on a spec sheet. But you need to take a peak beneath the spec sheet. If total spend is being driven primarily by consumption patterns and utilization rather than unit price, then optimizing procurement around the lowest sticker price is optimizing for the wrong variable.
Two organizations can pay the exact same per-token rate and end up with wildly different monthly bills. If one of them runs workloads on statically provisioned infrastructure that sits idle outside of peak hours, while the other is dynamically matching workload to available capacity as demand fluctuates, the rate card will look nearly identical but ultimately their bills won’t. This is part of why on-demand GPU rental rates have compressed so sharply over the past year. H100 pricing across major providers has dropped into a fairly tight band, with spot pricing dipping well below that during off-peak windows for workloads that can tolerate the flexibility. And yet, enterprise AI budgets have kept climbing regardless. This unit price compression and the total bill rising are part of the same consumption-driven story.
There’s also the matter of governance. Gartner's research errs by not fully diving into this problem in terms of procurement. If cheap, commoditized tokens make agentic capability more accessible, and organizations respond by deploying more agents for more autonomous work, then the real issue becomes one of visibility in near-real time, not before the bill arrives. This is something you won’t be able to do via vendor negotiations. Instead, you need to focus on infrastructure and orchestration, which most organizations are simply not currently equipped to solve because many of them can't track compute costs even in their current, comparatively simpler deployments.
Where you will find real GPU savings
Before we talk about where you can find real savings, let’s acknowledge that falling token costs are not absolutely meaningless. It’s a real development, and it will ultimately help reshape what’s economically viable to build.
What really matters, though, is tending to what happens between pricing and usage. In other words, intelligent orchestration. With it, workloads can be routed to the right infrastructure in real time, with dynamic price tracking across providers, and dispatching batch-tolerant work to wherever capacity is actually available and cheapest at that moment. This has been shown to yield cost reductions in the 40 to 60% range for workloads that can tolerate some scheduling flexibility.
Quantization, or the practice of reducing a model's numerical precision from something like FP16 down to INT8 or INT4, cuts memory footprint by two to four times and reduces inference cost by roughly half, while typically preserving well over 90% of the original model's accuracy on most tasks.
None of the above is speculative. These technologies are available right now and, indeed, they matter more to an organization's bill than any token price forecasting across the near and distant future.
The truth is that falling token costs and rising total spends are both real. If your only strategy is to wait for cheaper tokens, that’s not a real cost efficiency strategy. Ultimately, your team needs to know exactly what you’re running, where idle capacity is hiding, and then how to route workloads dynamically to that infrastructure that fits them. Only then will you achieve an AI budget that truly scales sustainably.
VOLT delivers GPU visibility, orchestration, and dynamic pricing
Engineered for AI startups, enterprise customers, and researchers, VOLT combines visibility with intelligent orchestration. Our dynamic pricing ensures you only ever pay for the compute you actually use, eliminating wasted budget and optimizing performance.
If you’re ready to scale your AI workloads with cost-effective compute, explore VOLT’s distributed GPU marketplace today, and deploy high-performance GPUs on demand.