It has become fashionable to quote our token index as support for a bearish view on the AI trade. We have pushed back gently a few times, because nothing here is definitive and people should come to their own conclusions. It now seems worth saying a little more.
In June we clarified what our LLM Token Index measures: what API users in our sample actually paid per million tokens. A usage-weighted price. Not token volume, not total spend. We noted it can be read, loosely, as revealed willingness to pay for frontier intelligence.
We should have emphasized the conditions. The binding one is that the intelligence content of a token stays stable. Over a few weeks that is defensible. Across a couple of release cycles it clearly is not. While a token is simply the wrong unit for intelligence, measuring capability itself without bias is also extremely hard. [Btw we will struggle with this measurement problem all over the place as AI proliferates delivering non-market economic value.]
The upshot is that the message from June landed. Maybe a little too well, because it created a new misread: that a falling index is necessarily bearish for the AI trade, since if token prices fall, model-layer margins must follow.
Two things are being conflated.
The deflation in our sample is real, and it is recent. It dates from the end of May. The index peaked at $2.07 on May 28 and sits at $1.00 as of Sept 21, down 52%. But the same index rose 67% from January into that peak, and few read the rise as bullish for lab margins. It is not bearish now. A usage-weighted price moves with the mix, in both directions.
What the mix actually says: more work, at least within our sample, is being routed to cheap, fast models, and labs keep shipping more mid-tier variants. Most everyday tasks never needed a frontier model. That is partial equilibrium for the users we track, and our methodology note is explicit that this index alone cannot separate substitution from efficient agentic routing.
Compute demand is a different question entirely, answered by different data series. Our H200 non-hyperscaler rental index has rerated through the summer: $2.87 average in June, $3.29 now, with a record $3.32 on Sept 19. B200 is $5.76, up 7% over the same stretch and 31% year to date. B300, our newest and thinnest series, is up 44% since inception in late April. H100 is off 7% from its August high, consistent with workloads migrating up the stack. Rents on the parts that are actually scarce are not signaling a demand stall.
A world of mass agentic use is one where cheap tokens are nearly all of the count. Total token usage should keep growing far faster than the price is falling, with most of that growth coming from cheap, fast, and increasingly open models. The usage-weighted price can keep falling anyway. Our view, not a finding from this index: the highest value-add work still routes through frontier models, and that is where most of the economics will accrue. In any event, the implication for compute is more, not less.
An astronomical number of tokens, most of them from cheap flash models, is what economy-wide AI proliferation should look like. It is not a demand stall.
Our LLM Token Expenditure Index should really have been named the “Token Expenditure Price Index” bc it’s an expenditure or usage-weighted average token price index. It tells you how much currently the entire market AI is paying for a million LLM tokens irrespective of models.
The naming might’ve led to some misinterpretations as some seem to have interpreted the index as either the total volume of token used or the average price of tokens. In reality, the index captures something more subtle than either interpretation: it tells us the marginal willingness to pay for LLM models.
Over the course of the year, while model token prices haven’t moved that much, the usage patterns have moved dramatically leading to the token index movement down and then up sharply as AI users moved en masse into using cheap open weight models and then en masse to the much more expensive frontier closed source models. From consumers to enterprises, everyone is Claude-maxxing!
More recently, as can be seen in the chart below, the token index has stagnated, which suggests that usage migration towards frontier models has slowed. Time will tell whether this is just a pause or an inflection in the trend as users move back towards open weights models. In a sense our token index could be roughly interpreted as a “quality premium” of frontier models over the much cheaper open source models (if we assume users and prices are both “rational”).
For more details on what we offer beyond the few indices we’ve listed on the Bloomberg Terminal, check us out at
silicondata.com and give us a holler! 😊