Last week marked an important narrative shift in AI. For the first time this year, the AI signals I was picking up from interviews, podcasts, and company news began to connect around a higher-probability version of something I raised in my weekly video two weeks ago: the risk of an AI capex air pocket.
That does not mean the AI trade is over, and it certainly does not mean there is an oversupply of compute. My core view remains that compute demand will stay structurally above supply for years. Training, post-training, inference, reasoning models, coding agents, enterprise agents, synthetic data, multimodal generation, private AI deployments, and eventually consumer agents are all moving higher in the years to come from a very low level. The more important point is that the demand curve is becoming less synchronized. The AI cycle is not ending. It is splitting.
That is how I read the recent flurry of news around Meta’s town hall and the reports that the company may begin selling excess AI compute capacity through a cloud business. At the town hall, Mark Zuckerberg reportedly acknowledged that Meta’s agentic AI progress had been slower than expected, saying the trajectory over the prior several months had not accelerated the way the company had hoped. Around the same time, Reuters reported that Meta was building a cloud business to sell excess AI computing capacity, based on an earlier Bloomberg report.
Those two stories should be read carefully. They are not evidence that AI compute demand is disappearing. They are evidence that the path from capex to monetization may be less linear than the market assumed which by itself will scare investors especially now that we are in the second derivative of earnings revisions for the semiconductors. Even in a world where aggregate compute demand remains above supply, the market can still experience an air pocket if the physical infrastructure curve runs ahead of the application monetization curve.
An air pocket is a timing and mix problem. It occurs when investors price a synchronized AI adoption cycle, while the actual demand arrives in separate waves. That appears to be the risk now. The market has been treating AI compute demand as one large category. More AI equals more tokens, more tokens equal more GPUs, and more GPUs equal more data centers, more power, more networking, more memory, more cooling, and more capex. Over the long run, that is still the base case for me. In the medium term, however, the demand curve is becoming more complicated.
There are at least three different AI demand curves now: coding inference, enterprise AI inference, and consumer personal-assistant inference. They are not moving at the same speed. Coding inference is accelerating. Enterprise AI is real, but it is moving with privacy, governance, data-control, and sovereignty constraints and now news stories about token price declines and competition from open source models. Consumer personal assistants are proving harder than expected because they require trust, memory, permissioning, payments, identity, and real-world execution. That split is the source of the air pocket risk.
The easiest part of the AI story for the market to understand has been the physical constraint. The industry needed GPUs. Then it needed more power, more networking, more memory, more data center capacity, more cooling, and more capital. Those constraints were visible and measurable. Investors could track the orders, capex budgets, lead times, and supply shortages.
The constraint is now migrating up the stack. For consumer AI, the bottleneck is not only whether the industry can build more compute. The question is whether the software layer can turn that compute into trusted, useful, autonomous products that people allow into their lives. The physical layer can be built faster than the trust layer.
That is why Meta matters. Meta is not a small company waiting for access to GPUs. It has capital, engineers, distribution, consumer data, social apps, open-source AI experience, and one of the largest AI infrastructure ambitions in the world. It spent tens of billions of dollar on AI talent a year ago. If Meta is acknowledging that agentic AI has not accelerated as expected, the issue is probably bigger than Meta-specific execution. It is a signal that consumer personal AI is harder than the market wanted to believe.
The evidence is now broad enough that this is no longer only a Meta issue. Apple continues to struggle to turn Siri into a true personal AI layer, including delays to the more capable, context-aware Siri features that were originally central to Apple Intelligence. Amazon has spent years trying to make Alexa more useful, yet the product still has not become the autonomous household operating system many expected. Google has enormous advantages in search, Android, Gmail, Maps, Calendar, YouTube, and Workspace, but it has not converted those assets into a dominant personal AI assistant that runs a user’s life. Microsoft’s Copilot, despite massive distribution inside Office and Windows, has also had a more uneven enterprise rollout than the original hype implied, with adoption constrained by workflow fit, governance, security, and the need for human oversight.
The common thread is that the first wave of personal and productivity assistants has not crossed the threshold from helpful interface to trusted autonomous operator. These companies still have enormous AI ambition and enormous compute needs, but the near-term consumer assistant layer is not ready to absorb the full infrastructure buildout. That is pushing each company toward other AI monetization paths, including cloud capacity, enterprise tools, developer agents, private models, search, advertising, devices, and data platforms, while they wait for the personal AI market to mature.
This is the key point. The AI race is becoming more complicated because the path to monetization is splitting. Anthropic can keep releasing more capable models and make rapid progress toward code-driven agentic AI, but that does not automatically mean the industry is making equal progress toward fully capable personal assistants. Companies that started with consumer assistants are being forced to find other ways to justify the infrastructure. Companies that started with enterprise workflows and coding agents are finding a more immediate path to measurable value.
The distinction between coding inference and consumer inference is central. Coding agents operate in machine-native environments. They can read a repository, edit files, run tests, inspect errors, call tools, search documentation, revise code, and try again. The environment gives them immediate feedback. The code compiles or it does not. The test passes or it fails. The pull request is accepted or rejected. There is a natural human-in-the-loop review process, supported by logs, tests, branches, diffs, and rollback mechanisms.
That makes coding an almost ideal environment for agentic AI. The task is difficult, but the environment is structured. The agent can act, observe, correct, and repeat. One human request can turn into dozens or hundreds of model calls as the agent searches files, debugs, runs tests, and improves the output. That is extremely compute-intensive, but it is also measurable. The enterprise can see whether the work was completed, whether the code shipped, whether productivity improved, and whether engineering throughput increased.
Consumer personal assistants operate in a very different environment. A true personal AI assistant has to do more than answer questions. It has to take action across messy human systems. It must understand ambiguous intent, remember personal preferences, interact with calendars, email, messages, shopping apps, travel systems, payment rails, subscriptions, healthcare portals, and other people. It must know when to act, when to ask permission, when to stop, and when to escalate.
There is no compiler for life. If a coding agent breaks a test, the test fails. If a personal assistant books the wrong flight, sends the wrong message, buys the wrong product, exposes private information, or misreads a social situation, the failure is more serious. That is why the consumer personal-assistant ramp is slower. Autonomy requires reliability, trust, identity, memory, permissioning, privacy, payments, and judgment.
This is where the market may have gotten ahead of itself. The most aggressive AI capex assumptions implicitly price a world of persistent consumer agents. In that world, every user has multiple AI agents operating in the background across communication, shopping, travel, finance, health, work, household management, content, and search. That world would create a massive inference shock because the unit of demand would change from one prompt and one answer to one goal and many model calls.
If consumer agents remain trapped in the assistant stage rather than the autonomous-action stage, that demand wave gets pushed out. Compute demand can still remain structurally strong because coding and enterprise AI can consume enormous amounts of compute. Frontier labs will continue to train and post-train more capable models. Enterprises will automate more workflows. Developers will use coding agents more heavily. Multimodal models will expand. The key issue is that the consumer inference wave may not arrive at the same time as the enterprise and coding waves. That timing mismatch creates the air pocket relative to expectations.
The other complication is that enterprise AI is not simply moving toward generic public-model token consumption. The enterprise market may increasingly move toward control. That is why recent comments from Alex Karp of Palantir and Databricks leadership matter. Karp has been making a broader argument that enterprises are frustrated with paying for tokens that do not create enough value, while also worrying about where their data, intellectual property, and institutional knowledge are going. Whether one agrees with Karp’s framing or not, the market signal is important. Enterprise AI customers are asking who owns the data, who controls the context, where inference runs, what happens to proprietary workflows, whether knowledge can leak, whether the system can be audited, and whether it can be governed safely.
Databricks is making a related point from another angle. CEO Ali Ghodsi has argued that AI does not mainly have an intelligence problem; it has a context problem. Enterprises need AI systems connected to proprietary data, governed workflows, and operational systems. That supports the same conclusion: enterprise AI demand is real, but it may not show up only as simple token consumption through a handful of public APIs. It may show up as private models, open-weight deployments, fine-tuned domain models, governed data platforms, hybrid inference, model routing, secure orchestration, and enterprise-owned context layers.
This still consumes compute, potentially a lot of it, but it changes the cadence and the winners. Some enterprises may continue using frontier APIs for the hardest reasoning tasks. Others may run smaller models privately for cheaper, safer, high-volume workflows. Some may fine-tune open models. Some may own the data and context layer while still routing specific tasks to external models. Some may insist on on-premise or private-cloud deployment for sensitive workloads. The key is that enterprise demand is becoming more fragmented, more governed, and more sovereignty-driven.
That matters for capex because as we can see in the increased volatility, many in the market may have assumed a cleaner path. Hyperscalers and AI labs would build massive centralized infrastructure, applications would plug in, tokens would explode, and monetization would follow. The reality may be messier. Enterprise demand can grow while also bringing more private deployments, more routing, more pricing pressure, more governance requirements, and more scrutiny over the value of each token.
That is why compute can remain undersupplied while the market still suffers an air pocket. The air pocket comes from the gap between aggregate demand and expected demand. If investors expected consumer personal agents, enterprise AI, and coding agents to accelerate together, but only coding agents accelerate immediately, enterprise AI moves more carefully, and consumer assistants take longer, the capex cycle faces a timing problem.
Meta’s possible cloud strategy fits into that framework. It may be an AWS-style move to monetize internal infrastructure. It may be a rational way to improve returns on capital. It may be investor messaging after capex concerns. It may also indicate that some internal consumer AI demand is not ready to absorb the infrastructure being built. All of those interpretations lead to the same market question: how quickly can AI infrastructure be converted into high-return AI revenue?
That is the question investors now have to ask across the sector. The bullish answer is that this is a short digestion phase. Consumer agents improve over the next six to twelve months, coding agents continue to scale, enterprise AI adoption broadens, and private AI deployments create another compute sink. In that case, the capex air pocket is shallow and compute demand quickly reaccelerates. The bearish answer is that consumer agents remain stuck for several years, enterprise AI becomes more selective and cost-disciplined, and capacity resale increases across the industry. In that case, compute demand still grows, but the market has to reprice the timing of returns on AI capex.
The most likely answer may be between those two extremes. Compute demand remains structurally strong, but the market becomes more discerning. It stops rewarding AI capex by default and starts asking where the demand is coming from, what kind of inference is being used, who owns the data, whether the use case is measurable, and how quickly the spend turns into revenue.
That is the next phase of the AI cycle. The first phase was about scarcity. The second phase is about digestion. The third phase will be about segmentation. Coding agents are the cleanest demand curve because they operate in machine-native environments with feedback loops and measurable ROI. Enterprise AI is the next major demand curve, but it is governed by privacy, context, sovereignty, and workflow integration. Consumer personal assistants are the largest eventual market, but they are delayed by trust, permissioning, payments, and real-world execution.
This is not a bearish argument but a more precise AI infrastructure reality. The easiest part of the trade is over. The AI mid-cycle slowdown is here and anyone riding the AI train looking for no bumps, are learning, extrapolation can be dangerous if you see everyone else on the train too. Compute demand will remain above supply over time, but the market can still face an air pocket if the next wave of demand arrives later, differently, or in a more fragmented way than investors expected. The AI cycle is not ending. It is splitting. The next investment question is no longer simply who can build the most compute. It is who can match that compute to the demand curves that are real today, while surviving the gap before the largest consumer-agent wave finally arrives.