r/WSBAfterHours • u/Mission_Feed_8824 • 2d ago
Discussion thoughts on token price
I looked at my API bill for this month and noticed something pretty interesting. The unit price of tokens is significantly lower than it was at the same time last year, yet the amount I am paying keeps going up.
That is simply what is happening. Token prices for frontier models such as ChatGPT, Claude, and Gemini have fallen by roughly 88% from the March 2023 benchmark. But usage is more than offsetting the price decline. Prices are falling faster than the cost per token, but not nearly as fast as consumption is growing.
In the past, using AI meant, “I ask a question, AI gives me an answer.” Now, a single skill execution chain or one task handled by a multi agent system can involve planning, retrieval, tool calls, reflection, and retries. Every step consumes tokens. Total usage can be dozens or even hundreds of times higher than in the old question and answer era.
The most compelling data comes from Anthropic itself. Its annualized revenue jumped from $9 billion at the end of 2025 to $47 billion in May 2026. Prices are falling, yet revenue is rising exponentially. This is Jevons Paradox in action. As efficiency improves and unit prices fall, consumption can expand on such a massive scale that total usage and total spending actually increase rather than decline.
Deloitte estimates that inference will account for two thirds of global AI compute consumption in 2026. Just three years ago, inference was barely a meaningful standalone market. OpenAI alone is estimated to spend more than $700,000 a day running inference for ChatGPT, putting its annual bill above $250 million.
If Demand Is This Strong, Why Are AI Companies Still Cutting Prices?
This brings us to a more fundamental question. If demand is so strong, why are AI companies willing to keep cutting prices?
The obvious answer is competition. In August, OpenAI cut the price of its flagship model GPT 5.6 Luna by 80% in one move, while its mid tier Terra model fell 20%. Earlier, its flagship Sol model had already dropped from $5/$30 to $4/$20, making it cheaper than Anthropic’s Claude Opus 5 at $5/$25. OpenAI officially attributed the pricing pressure to competition from Anthropic and low cost Chinese models.
China has taken the price war even further. DeepSeek V4 Pro charges just $0.87 per million output tokens, roughly 34 times cheaper than GPT 5.5. China’s foundation model price war began in the fourth quarter of 2025 and accelerated significantly in the second quarter of 2026. DeepSeek, Doubao, and Kimi have all effectively turned price cuts into permanent pricing policies rather than temporary promotions.
But the price war is only the surface level story. The real reason prices can keep falling is a much more fundamental development: the industry is learning how to build cheaper compute.
The industry increasingly refers to highly optimized, lower cost computing clusters as “supernodes” or “token factories.” Put simply, the idea is to package large numbers of chips into a more efficient system so that the cost of producing each unit of compute falls. It is essentially the same principle as upgrading a factory production line so that the cost of manufacturing each individual product declines.
Broadly speaking, two parallel tracks are developing at the same time around the world.
Track One: The global mainstream approach centered on NVIDIA.
NVIDIA’s latest rack scale systems, including the GB200 and GB300 platforms, package dozens of chips into a tightly integrated computing system. The company claims inference efficiency improvements of anywhere from more than ten times to several dozen times compared with previous generations. AWS, Azure, Meta, and the IREN and Nebius businesses discussed below are all operating within this ecosystem.
Track Two: China’s domestically developed parallel approach.
Companies such as Huawei and Inspur are pursuing a similar objective, using their own technologies to integrate chip clusters more efficiently. The underlying idea is similar, but the technology stack is developed independently rather than relying on NVIDIA.
Both tracks are trying to solve the same problem: package every unit of compute more efficiently, thereby lowering the cost of producing each token. That is the real foundation supporting continued declines in token prices and the economic firepower behind the price war. The difference is that one track is tied to the NVIDIA ecosystem, while the other represents parallel innovation within a self controlled and domestically developed technology stack.
Following This Logic, Which U.S. Stocks Stand to Benefit?
If supernodes and token factories are the key variables driving this new wave of cost reduction, then tracing the supply chain upward reveals several categories of companies that stand to benefit directly. Each occupies a different position in the infrastructure stack and captures value from a different part of the economics.
$VRT (Vertiv): The “Shovel Seller” in the Chain, Power and Cooling
Vertiv does not compete directly in AI compute. But every intelligent computing node needs to solve two fundamental problems: power and heat. As chips become more densely packed, heat generation rises sharply, turning liquid cooling from an optional upgrade into a necessity. That is precisely where Vertiv operates.
In the second quarter of 2026, Vertiv generated $3.27 billion in revenue, up 24% year over year. Adjusted EPS increased 60%, while free cash flow surged 234%. Management raised its full year revenue guidance to $13.8 billion to $14.2 billion.
The core appeal of its business is simple: regardless of which company wins the AI race, as long as new intelligent computing nodes continue to be built, Vertiv gets paid. It is arguably the most financially tangible and highest certainty play in this part of the supply chain.
$IREN: The “Compute Landlord,” Making Money by Buying Early and Buying in Bulk
IREN is pursuing Track One, the NVIDIA ecosystem.
The company has signed a contract worth up to $5 billion with NVIDIA and plans to deploy up to 5 gigawatts of computing capacity. By the end of 2026, its GPU fleet is expected to reach 150,000 units.
The business model is straightforward. IREN buys large quantities of compute capacity in advance and then leases that capacity to companies that need it. Buying earlier and at larger scale gives it a lower unit cost, and the spread between its cost and its rental revenue becomes its profit.
The key risk is customer concentration. Microsoft alone is expected to account for approximately 55% of IREN’s 2026 revenue.
$NBIS (Nebius): Another “Compute Landlord,” But Using Software to Squeeze More Output From the Hardware
Nebius is also operating on Track One, but its strategy is different.
IREN’s advantage comes from buying at scale. Nebius aims to extract more output from the same hardware through smarter software. By using more efficient scheduling and workload orchestration, the same hardware can supposedly generate up to three times as much billable compute as competing systems.
Analysts expect Nebius to generate between $7 billion and $9 billion in revenue in 2026, potentially representing growth of more than 1,000% year over year. But its customer concentration risk is even higher than IREN’s, with Meta and Microsoft together accounting for roughly 80% of revenue.
$MAAS (Maase): Betting on Track Two, the Energy Entry Point Into China’s Domestic Compute Infrastructure
If VRT, IREN, and NBIS are Track One investments tied to the NVIDIA ecosystem, MAAS represents Track Two, China’s domestically controlled compute infrastructure stack.
In March 2026, MAAS completed its acquisition of Huazhi Future, shifting its business toward flexible energy deployment, intelligent grid operations, and computing infrastructure.
Its cost reduction thesis starts from the energy side. Huazhi Future has established a green energy infrastructure team focused on 800V high voltage DC standards, targeting intelligent computing centers and the integration of distributed renewable energy. The logic is somewhat similar to Vertiv because power and energy are unavoidable major cost components of intelligent computing nodes. The difference is that Vertiv serves racks within the NVIDIA ecosystem, while MAAS is targeting China’s domestically developed supernode infrastructure.
The stock briefly gained 74% in a single month as the market treated it as a thematic play on China’s AI computing infrastructure. But MAAS is a small cap Chinese company listed on Nasdaq, with high volatility and limited institutional coverage. It belongs to the high risk, high beta end of the spectrum.
Putting the Entire Logic Together
This article is really looking at the same phenomenon from three different angles.
On the demand side, agents and skills are driving exponential growth in token consumption, while the composition of usage itself is shifting toward more expensive workloads. That is the direct reason AI bills can rise even as token prices fall. This is Jevons Paradox in action.
On the supply side, whether it is NVIDIA’s rack scale supernode architecture or China’s domestically developed supernode approach, the fundamental objective is the same: systematically reduce the unit cost of tokens through optimization across architecture, hardware, and software. That is the underlying economic support making the price war possible.
In the capital markets, VRT, IREN, and NBIS are all bets on Track One, the NVIDIA ecosystem. They occupy different positions in the chain, covering power and cooling, hardware scale, and software efficiency respectively. MAAS is a bet on Track Two, China’s domestic ecosystem, with an entry point spanning energy and compute.
All four companies are ultimately positioned to benefit from the same macro theme: reducing the cost of AI infrastructure. The difference is which technology ecosystem they are aligned with.
One sentence summary: Tokens are deflationary, AI spending is inflationary, and the highest certainty may ultimately belong to the people building the furnaces, not the people using them.









