Summary and learnings from the Silicon Data webinar: GPU & Token Pricing
Yesterday, we had the pleasure of hosting Silicon Data founder and CEO Carmen Li, along with Head of Research Steve Hou. This call added depth and clarity to understanding the emerging AI token and GPU pricing signals. In recent months we have been clear that price discovery and the ability to interpret those prices will be critical for investors. (Emerging Price Discovery for Tokens and Compute, 6/9/26). Silicon Data has developed the often-cited LLM Token Expenditure Index (Bloomberg: SDLLMTK) and is building a robust data platform around compute pricing — the foundation for tradable compute financial products.
Please see our key takeaways and some highlights and observations from the webinar.
Replay access HERE
Key Takeaways:
- Price discovery is advancing, but early. Observable, normalized benchmarks now exist for both tokens and GPU rentals, with CME-listed derivative products a few months out — a further step toward reading the signals from a rapidly changing AI market.
- The direction of the Token index chart does not indicate weaker AI adoption; rather, a move higher or lower indicates the impetus to move into higher or lower priced models or implies downward pressure on AI token prices. In recent months, not only have lower priced Chinese models become more prevalent but also Meta and Grok are releasing higher-performing models at significantly lower price points than Anthropic or OpenAI. Silicon Data’s head of research noted that the volume trend of AI usage — tokens or otherwise — is up and to the right.
- Spot and futures pricing for H100 and earlier GPU models continue to be firm, indicating a tight market for compute. The long-term pricing discount to spot has narrowed as well, which further highlights the state of the market.
- Silicon Data’s CEO believes the compute financial-products market could become one of the larger financially traded instruments behind oil and fuel commodities. The company is developing robust data sets and index products to accommodate this expansion.
- Financial products around GPU’s appear to be reasonably complex. Factors that need to be taken into account include the type of GPU, the performance benchmarking and end of life value of the GPU. As more generations of compute come into the market, the complexity of these products is likely to increase.
- International dynamics are increasingly worth watching, and Silicon Data has the global (ex-China) data to track these regional price ranges.
Token Pricing and the LLM Index
- The LLM Token Expenditure Index (SDLLMTK) is an expenditure-weighted price index across hundreds of models, PCE-like in that it can reflect model substitution and token pricing changes.
- Recent moderation is a mix shift toward cheaper open-weight models after end user cost scrutiny — not a demand signal.
- Total token volume and total expenditure are both rising monotonically; lower prices are driving adoption (Jevons paradox). It is the slope that varies, not the direction.
- The index reflects the price-sensitive corner of the market (public inference platforms) can be a leading indicator of enterprise behavior; it is not a full-market view (no privileged hyperscaler/frontier-lab data).

Source: Silicon Data
GPU pricing and futures
- Emerging financial products around compute are comparable in some ways to the development of oil futures markets. Later this year, the first batch of CME listed products will start to trade. Cash-settled futures and options on GPU indices — H100, A100, and more.
- These products will create risk management mechanisms to clear markets between buyers and sellers of compute.
- Physically delivered futures create complexity around grading and benchmarking. Silicon Data has developed Silicon Mark. Like a “Carfax for GPUs”: immutable, timestamped performance verification, that can be used by banks and insurers to underwrite residual-value contracts
Signals and trends
- Recent high priced SpaceX compute deals paved the way for Meta to similarly consider renting out available compute. Silicon Data does not see this as a wave of supply — notably, compute pricing remains robust — but rather as opportunistic. These deals equated to over $15/hour, well above market rates.
- H100 is the market’s summary statistic: on-demand softened recently, but 1-year forward/reserve prices rose from late June into July, coincident with reports of Meta considering leasing capacity — providers raising, not discounting.
- Term structure has moved up wholesale since November and flattened from backwardation toward contango, signaling providers are comfortable not discounting long-term contracts and prefer rolling short-term to retain pricing flexibility.

Source: Silicon Data
- Hyperscaler vs neocloud GPU rate spread: H100 at ~$2.50/hr (neocloud) vs ~$7.24/hr (hyperscaler), the premium reflecting bundled software/services. While the pricing spread is wide, it is seeing some convergence as neo-cloud pricing has been moving higher.
- International: meaningful cross-sectional pricing variation across North America, Europe, APAC, and the Nordics; adoption still early ex-US (Europe/South America lighter than hoped), with China on a separate supply track.
- Utilization remains difficult to observe; no data product yet, though Silicon Data is working toward more visibility.
Please reach out with any questions or follow up on this topic.