GLM 5.3 now online!Try it →
TokenGO

─── BLOG

Compute prices could surge over 10x in the future

Aug 20, 2026 · TokenGO

Revenue is growing far faster than compute supply. If that gap holds, AI compute gets dramatically more expensive and winner-takes-all dynamics harden.

This is a calm deduction of compute pricing in the AI arms race. The core logic is crystal clear: revenue growth far outpaces supply growth, meaning compute will become increasingly expensive.

The winner-takes-all dynamic will likely solidify into an iron curtain. Those trying to catch up must figure out exactly how to pay the astronomical entry fee.

A set of data worth watching

Let us look at a few numbers first.

Anthropic's revenue has grown 10 times year over year. By the end of this year, this figure is highly likely to reach $100 billion to $150 billion.

The spot price for AI compute has risen by over 40% since its low point in February.

Google and Anthropic are renting 110,000 GPUs from SpaceX for $900 million a month, which is roughly double the spot price.

Anthropic revenue, spot prices, and a SpaceX rental deal

If Anthropic's growth trend continues, its revenue will hit $1 trillion by the end of next year. While this trend could easily halt, assuming it actually continues, what would that world look like?

The key lies in a simple mathematical relationship. Revenue grows by 10 times, but compute supply only grows by 3 times. How do we fill that gap?

Revenue grows 10x while supply grows 3x

Three options, one outcome

Lab compute grows by about 3 times annually. To achieve a 10x revenue growth with only a 3x compute growth, some combination of the following three scenarios must occur:

  1. Lab profit margins increase.
  2. Compute prices rise significantly.
  3. Labs allocate a larger proportion of compute to inference.

Right now, all three are happening.

Margins up, prices up, or more compute on inference

Anthropic's profit margin has risen from 40% in 2025 to a level that could exceed 80% for its Fable inference service this year. Meanwhile, compute spot prices are up over 40% compared to the February trough. Finally, OpenAI spent about a quarter of its compute budget on inference in 2024, and now that ratio is close to 50% or even higher.

However, labs are highly reluctant to pursue the third option of shifting an ever-growing share of compute to inference.

A straightforward joke explains this perfectly. The whole point of inference revenue is to convince investors to give you more money so you can buy more compute to train larger models. If you spend most of your compute on inference, you are basically declaring that AI progress has stalled, and your business has essentially turned into a cloud service provider.

Therefore, the two remaining options are that profit margins continue to rise or compute becomes vastly more expensive.

How high can profit margins go?

If one or two leading labs are clearly ahead of their competitors, profit margins can indeed go higher. Your profit margin depends entirely on how much better you are than the second-best alternative.

For the margin effect to dominate, profit margins would need to reach around 95% by the end of next year. This sounds crazy. However, the fact that AI lab revenues continue to grow at astonishing rates is highly unusual in itself.

Thus, only one effect remains to explain how the "$1 trillion revenue by the end of next year" scenario plays out, which is that compute becomes incredibly expensive.

Compute prices are already rising

When we focus on the specific type of compute that labs actually need, the price increase is even more pronounced.

Labs clearly cannot rely on spot instances. They need to ensure the security of model weights and customer information, and they need sufficient scale to achieve optimal utilization and flexibility.

A realistic reference point is that Google and Anthropic rent compute from SpaceX, paying $900 million monthly for 110,000 GPUs (a mix of GB200s and GB300s). This is roughly double the hourly spot price for these GPUs. The spot price itself is already 40% higher than it was in February.

The crucial deduction: the $250,000 H100

Imagine this scenario.

If a truly human-level software engineer could run on the compute equivalent of a single H100, based on current market rates for software engineers, the annual rental cost for one H100 should exceed $250,000.

This is 15 times the current spot price.

One H100 priced like a human software engineer

You might argue that if AI suddenly adds 10 million software engineers, the marginal value of a software engineer will drop, meaning an H100 might fail to generate 15 times its current revenue.

However, this logic is flawed. Applying it to humans turns it into the classic "lump of labor fallacy." Economists generally agree that high-skilled immigration brings distinct benefits through specialization and innovation rather than depressing the wages of existing high-skilled workers.

Perhaps the scale and speed of this labor supply shock are so massive that this general rule of thumb stops applying. Yet, if we believe standard economic theories regarding labor, then the marginal value of compute, and consequently its marginal price, could become staggeringly high.

What will this lead to?

First, catching up becomes incredibly difficult. If software engineering is automated by 2028 and compute prices are 15 times higher than today, latecomers without revenue will find it virtually impossible to compete with frontier labs for compute.

Second, the winner-takes-all dynamic will become heavily entrenched. If you have the best model, you can charge a vastly higher profit margin than you do now. This is the Alchian and Allen effect in economics. If the price per H100 hour is $20 instead of $2, running a dumber model on this compute is simply foolish. Since you are paying such an exorbitant fee for the underlying hardware, you are better off using the absolute best and most efficient model directly.

Third, many current AI applications will become obsolete. Part of the reason AI is relatively cheap today is that it completely lacks the ability to do many things top humans can do. At a certain point, this will change. Using GPUs to churn out short-form video spam will gradually die out because the compute will simply be too expensive.

Harder catch-up, entrenched winners, and obsolete cheap apps

Could this prediction be wrong?

It is certainly possible for this prediction to be wrong.

History is full of incorrect predictions about scarcity. The wager between Julian Simon and Paul Ehrlich is a classic example where Ehrlich bet that the price of a basket of commodities would rise before 1990 instead of falling, and he lost. Market signals and human ingenuity consistently find more efficient ways to utilize scarce resources.

I lean toward the belief that Simon and Ehrlich's basket of commodities is the wrong reference category for compute.

The elasticity of compute supply is far lower than the mining of various metals, and its ability to absorb massive demand shocks is significantly weaker. The 3x annual growth in compute comes from the product of the following factors:

1.4x Moore's × 1.2x fabs × 1.8x wafer share ≈ 3x supply

A longer-term perspective

At some point in the future, compute will become cheap again.

When robots can transform coastal silica sand and copper ore into computers, the price of compute will drop close to the cost of raw materials and tools.

However, I am only discussing the current phase, where AI compute grows at a mere 3x per year. This rate is wholly insufficient to offset the price effects driven by the massive annual improvements in AI utility.

Until then, compute will go through a period of sky-high prices and winner-takes-all dynamics. For labs without massive revenue backing, this is a challenge that demands serious thought.

─── NEXT

Try TokenGo's serverless inference API now →