In the past week I have had the opportunity to connect with a handful of operators and developers across the AI factory build-out, as well as several of the industry participants leading the way on price discovery for tokens and GPUs.
The overarching theme is that the market remains compute constrained for the foreseeable future and that hyperscalers and AI builders are likely to continue front loading the AI build-out. There were no indications of a pause from the conversations. The underlying reasoning stems not only from increased demand and new product development, but also from execution challenges slowing the pace of getting compute online. AI factories are relatively new and evolving and require significant learning along the way. The beneficiaries will be the top executors, from infrastructure players to bare metal powered shell and neo-cloud operators.
Demand continues to outrun supply, and the point at which the two converge keeps moving further out — sustaining a tight market and, with it, pricing power for those who can execute..
Below are some key takeaways and investment considerations.
Demand for compute remains strong: Hyperscalers are building capacity now and show no sign of slowing spending plans. Industry players echo the need for hyperscalers to build capacity to compete as well as emerging venture backed startups and enterprises building products and workflows. There is a continuing shortage of compute, which in turn is holding back the next wave of more powerful models and products that would otherwise come to market — the constraint is supply, not appetite.
The price of long-term compute contracts appears to be heading higher than the Q1 neo-cloud deals struck in the $10–12M/MW range. First, xAI raised the bar with their Colossus deals; secondly the increased power of compute (Vera Rubin) and rising costs of data center builds can drive the next deals higher. In addition, there is a wide gap between the price for long-term take or pay deals and spot market capacity. The middle ground is a logical place for new compute transactions to come together.
The value of the platform and capacity are important differentiators. Confirming the pricing-power dynamic above is AWS’s recent announcement to raise certain EC2 Capacity Block reservation prices by roughly 20%, effective July 1 — its second increase this year, following a ~15% hike in January — reflecting a premium to other providers for similar compute. This confirms that pricing power is accruing to those with capacity and a strong platform.
Labor and supply chain generate delays: Labor, in particular electricians and other skilled trades, remains tight. This, combined with supply chain timing issues, can create cascading delays for data center buildouts. A consistent message was that delays would become more common and discussed within the industry.
Orchestration of compute is an emerging bottleneck: Probing directly on what it takes to bring vast quantities of leading-edge compute — Vera Rubin, TPU clusters — online yielded responses that these are complex ecosystems to stand up with mixed results about effective utilization in early deployments. Hence why many are stating that not all GPU’s are created equal when placed in an AI factory. Success depends on aligning many moving parts, and the industry is still learning as it goes.
This drives a flight to quality: the operators who master the full chain of events— power delivery, frequency and load management, thermal management, chip-level orchestration — are the ones who will win and generate the maximum ROI. With costs for compute exceeding $50B, losing even 10% of a cluster to orchestration failure is a serious return impact. A reported example, Crusoe, which is at the frontier of bringing on early AI factories like Stargate, has faced many challenges to bring AI factories online (The Information, 6/8/26) underscoring this is not plug-and-play.
Power: easier to locate grid power after 2027: Power availability and grid connections are strained in the near term; however, locating additional power is less difficult in 2028 and beyond. It raises some questions as to where Behind the Meter (BTM) and Bring Your Own Generation (BYOG) projects and cumulative volumes start to converge with large scale grid capacity being brought online. The near-term constraint is a catalyst for smaller, modular power deployments that can avoid long interconnection queues, but the medium-longer term grid picture may start to look different as we approach 2030.
Price discovery is in early innings for Tokens and GPUs: We have spoken with Silicon Data (the firm behind the Bloomberg GPU/LLM benchmark) and corroborated with operators directly, building on our earlier work (Emerging Price Discovery for Tokens and Compute, 6/9/26). They are doing great work to forge this new path.
The reality is that only a small fraction of the market data is available; neo-clouds have many large take-or-pay contracts, so there is little incentive to publish real-time utilization or share specific frontier model usage splits. Directionally, an index like the Silicon LLM benchmark is getting the direction of travel (increasing use of lower price models in recent weeks) correct; however, this represents a sliver of the overall market today.
GPU pricing and futures are even earlier on the curve. Spot and limited depth futures pricing is only available for older model GPUs like the H100. Additionally, and more importantly – not all GPUs are created equal, and the host platform (such as AMZN’s Bedrock) can have large impacts as to the effectiveness of the compute. These markets and pricing mechanisms contemplated will be important for the market as it develops, however near-term volatility is not a strong signal either way for the health of the AI trade.
Orbital data centers — operators are skeptical: Everyone we spoke with is decidedly skeptical of orbital data centers (ODCs). The pushback is grounded in the operating reality: those who run GPUs at scale daily know how much hands-on intervention is required — swapping a cable, changing out a CPU, and countless other routine maintenance items. None of that is possible in space. From an operating perspective, any small maintenance item effectively puts millions of dollars at risk of loss in space, which is hard to underwrite. While skeptical, it was clear ODC’s will be watched and that news flow around testing and validation is being followed now that SpaceX is public.
NVIDIA remains dominant but increasingly complex to operate: GPU orchestration is not getting easier as the compute becomes denser and large-scale implementations run into the hundreds of thousands of GPUs. NVIDIA hardware remains the clear standard for large AI workloads, but the operational complexity of running at scale is increasing with each hardware generation. TPUs were noted as next up as a contender in AI compute (the first TPU neo-cloud is still coming to market) while Trainium continues to be viewed as most early stage and challenging. The implication for investors is that operational expertise — not access to hardware — may be the more durable moat at the neo-cloud operator and infrastructure layer.
Longevity of older GPUs: Early generation GPUs such as the H100 continue to show promise of longer life as evidenced by both current rental rates as well as market feedback. This message has been consistent year to date; GPU’s will be repurposed rather than retired at 5 years, with software and orchestration technologies built to harness them more effectively on less time-sensitive AI tasks. The implication of this life extension points to more greenfield building for leading edge compute (with liquid cooling and emerging 800V) versus operators taking the financial hit for downtime to retrofit and lose valuable sunk capex on HVAC equipment.
NIMBYism and the edge: Data center developers seem to be fielding increasing moratorium and data center opposition questions. It no longer is popular to advertise large campus projects; however local community relationships are becoming key foundational steps. Blending these challenges with growing demand for closer edge AI compute, some developers are pursuing sub-50MW edge deployments. Depending on the market, these can avoid lengthy studies or reviews and satisfy a need for compute closer to the market. Physical AI (robotics and autonomous vehicles) requires proximity for low-latency operations creating a structural pull toward this model.
Investment considerations
- Neo-clouds: Expect differentiation via execution. Those that can orchestrate the compute and meet build-out timelines will be able to grow more quickly and receive higher prices for the compute.
- Power: Near-term constraints have driven a massive grid and BTM/BYOG investments. The supply-and-demand curve of this evolution is important to start modeling out in more detail — capital goods, power operators, and the type of BTM power to determine the demand heading into 2030.
- Price discovery: Token and GPU prices will be important to follow, but we would continue to caution that token pricing reflects only a fraction of the overall market and that GPU futures markets are in their infancy. Near term price volatility for H100 rental prices should not be market moving. Development of the compute futures markets with higher transaction volumes will be important to monitor heading into 2027.
- Liquid cooling: Infrastructure is decidedly moving toward liquid cooling, and with the ongoing build-out and investment, companies with this portfolio stand to benefit from a large TAM. NVIDIA’s recent release of 45°C operating temperatures (the Rubin DSX reference design) should further highlight potential disruption to traditional HVAC data center component share.
Please reach out if you’d like to discuss further.