Every time you ask your phone to rewrite an email, you are buying digital grain by the bushel. You do not purchase software licenses, compute cycles, or cloud storage. You buy tokens, the fundamental currency of modern artificial intelligence. A single token is roughly three-quarters of an English word or 3.5 bytes of data. That microscopic billing unit is quietly rewriting the economics of enterprise technology.
Think of an artificial intelligence model as a commercial freight truck delivering packages of words. The engine does two completely different kinds of work during a single delivery run. First comes prefill, where the processor ingests your entire query in one massive, parallel burst of raw calculation. Then comes decode, where the model outputs its response one syllable at a time.
Decode is where the hardware actually sweats. Generating text is not limited by sheer computing muscle, but by memory bandwidth. Each generated token relies on the previous one, forcing the system to stream massive model weights alongside a ballooning memory log called a key-value cache. Because that cache can expand to several times the size of the underlying model, the real bottleneck is how fast memory chips can ferry data across silicon.
Yet the cost of that journey is in freefall. Research from Epoch AI shows the price of equivalent model performance drops roughly 47% per quarter since 2023. That amounts to a 13-fold annual reduction, falling four times faster than genomic sequencing. GPT-3 reached high-school level benchmark accuracy at $60 per million tokens in late 2021. Today, small open models match that accuracy for six cents, an astonishing thousand-fold collapse in three years.
In traditional software, severe price collapses destroy margins and terrify equity investors. In artificial intelligence, deflation functions as jet fuel. When reasoning becomes essentially free, businesses consume it in torrents rather than sips. OpenAI showed that answering a doctoral-level science problem cost 30 cents on an early reasoning model, before falling to $0.0004 on a lightweight successor eighteen months later.
This collapse unlocks what Gartner calls the inference paradox. Token unit costs may fall 95% by 2030, yet the cost of running one agentic workflow climbs more than fivefold through 2028, because autonomous agents demand far more thinking steps to verify results. A complex autonomous agent does not produce a single answer. It plans, writes code, checks its mistakes, and queries other agents, consuming up to twenty times more tokens per task. As individual words get cheaper, total word consumption explodes.
Look at the cloud giants to see this dynamic compound in real time. Across its products, Google went from roughly 480 trillion tokens a month at its 2025 developer conference to over 3.2 quadrillion a month a year later. Its model APIs alone now run at about 22 billion tokens a minute, up from 16 billion the quarter before. Alphabet posted an 82% jump in cloud revenue alongside a $514 billion backlog because enterprise usage easily outstripped price declines.
Hardware architects are racing to own this cost curve directly. Amazon built its custom 3nm Trainium3 accelerator with 144 gigabytes of high-bandwidth memory specifically to solve the decode memory bottleneck. Nvidia engineered its Blackwell architecture to drop token generation costs up to tenfold compared to prior generations, cutting certain large model costs from 20 cents to five cents per million tokens. For these chipmakers, keeping older silicon rented demands outrunning their own deflation.
The bear case argues that this deflation is self-inflicted and will ultimately destroy the tollbooths. If token prices fall 13-fold annually while models become interchangeable commodities, inference turns into long-distance telecommunications where pricing races toward zero. Open-weight Chinese alternatives already offer cached input rates at $0.003 per million tokens, functioning as a brutal clearing price rather than a temporary promotion. Hardware techniques like quantization and sparse activation steadily drain the revenue a server rack generates per hour. The warning signal for equity holders will arrive when cloud token volumes keep compounding while cloud revenue growth abruptly decelerates.
For now, compute demand is outpacing unit deflation by a wide margin. Deloitte estimates inference workloads will account for roughly two-thirds of all AI compute, up from a third in 2023. The market does not reward developers for building precious models that hoard margin. It rewards the infrastructure landlords who can manufacture intelligence cheap enough to waste.
What matters next for the equity re-rating is whether enterprise workloads shift from human prompts to multi-agent loops. The number to watch is the gap between two curves: volume compounding, price per token falling. If agents multiply faster than chip designers cut input costs, the token economy keeps expanding.
The cheapest thinker wins the contract.
Analysis, not advice. We may own these names and we may sell them at any time.
AI does not sell software seats, it sells tokens, and the price of a token keeps collapsing while the number of tokens keeps exploding. That is why a 13-fold annual fall in the price of thinking has coincided with a sevenfold jump in Google's monthly token volume and a doubling of inference's share of AI compute. The test for the stocks is not whether prices keep falling, because they will. It is whether demand compounds faster, and whether the company collecting the toll also owns the cost curve underneath it.




