This weekend’s video highlighted an important reality now hitting the AI growth and adoption story: the economics of AI are increasingly being shaped by physical constraints. It was only six months ago that the number one fear for investors was around and AI bubble and overbuild. Now after the exponential rise in token usage, the issue is no longer just model quality, hallucinations or falling abstraction costs. Demand is rising faster than the underlying infrastructure can scale, and the supply side is running into real-world limits in compute, memory, and power. That matters because even when token prices look manageable in isolation, the broader cost of deploying AI at scale is being pushed higher by scarcity underneath the surface. The result is that enterprise AI, which has only recently begun the adoption process, is beginning to behave less like frictionless software and more like an industrial system whose economics are tied to physical bottlenecks.
The recent Uber story is a useful place to start because it captures the central tension in enterprise AI right now. Companies are adopting AI tools on the assumption that they will compress development cycles, automate portions of knowledge work, and eventually improve margins by replacing some labor. But the first real wave of enterprise usage is showing something more complicated. The productivity gains are real, yet the cost curve is rising faster than many management teams expected. That is why Uber CTO Praveen Neppalli Naga’s reported comment landed so hard: “I’m back to the drawing board because the budget I thought I would need is blown away already.” According to reporting on Uber’s rollout, the company pushed Claude Code broadly across engineering, saw adoption surge, and exhausted what it had expected to spend on AI far earlier than planned. That is not the story of a failed tool. It is the story of a useful tool that is starting to bill more like an engineer than a piece of software.
That distinction matters because the industry is increasingly selling AI not as software access but as labor substitution. Goldman Sachs, in a report highlighted by Business Insider, said companies are increasingly positioning AI workflows as a “unit of labor” or a “unit of productivity” rather than a traditional seat-based product. It is easy to see why vendors like that framing. Labor budgets are bigger than software budgets, and a product that can credibly claim to replace hours of work can command much larger deal sizes than one that merely adds convenience. But that shift in framing also changes the economics for the customer. Once AI is sold as labor, the customer is no longer buying a mostly fixed-cost SaaS subscription. The customer is buying a metered service whose usage rises with ambition. The more tasks the system handles, the more tokens, inference, orchestration, memory, and compute it consumes. What looks like software on the surface starts behaving like infrastructure underneath.
Aaron Levie has been unusually candid about what this looks like inside companies. In April he described “tokenmaxxing” as a real budgeting issue, noting that many enterprises lock in operating-expense budgets early in the year and are now having active internal fights over how to allocate compute. He also made a second point that is even more important for the margin debate: most companies are not yet talking about eliminating jobs because of agents. Instead, they are using AI to do things they previously could not prioritize, or to increase output rather than simply cut payroll. That sounds bullish for adoption, and it probably is. But it is not immediately bullish for margins. If AI expands the amount of work a company chooses to do before it materially shrinks headcount, then the enterprise ends up carrying both cost structures at once. Levie’s broader warning is that coding agents already consume orders of magnitude more tokens than earlier chat-style use cases, and that the same consumption pattern is likely coming for the rest of knowledge work. That is the real economic issue: agentic software creates variable costs that can scale much faster than traditional software buyers are used to managing.
The pricing architecture reinforces that point. Anthropic’s current public pricing shows Claude Opus 4.7 at $5 per million input tokens and $25 per million output tokens, while Claude Sonnet 4.6 is priced at $3 and $15 respectively. Those numbers may sound small in isolation, but they stop looking small when you multiply them across thousands of developers, long context windows, repeated retries, tool use, background agents, and persistent workflows running all day. This is not classic SaaS economics, where the marginal cost of one more workflow often rounds toward zero once the license is sold. It is usage-priced software layered directly onto compute. The more sophisticated the workflow, the less relevant seat count becomes and the more relevant token volume becomes. That is why enterprises are now having to think about model routing, spend controls, caching, and task design as financial disciplines rather than just engineering choices. AI is not only becoming smarter; it is becoming more expensive to use badly.
This is where the Wall Street Journal’s recent reporting becomes especially important. The Journal reported this month that the AI gold rush is rapidly drying up the supply of computing power, forcing companies to ration products and confront reliability problems as demand outruns available capacity. In a separate report, the Journal also described Oracle laying off workers while continuing to build out costly AI data centers. Put those two developments together and a broader picture emerges. AI costs are not rising solely because model companies have chosen aggressive pricing. They are rising because the entire infrastructure stack beneath AI is under pressure: compute is constrained, power is scarce, and the capex required to keep expanding supply is enormous. This matters for end users because scarcity at the infrastructure layer eventually shows up as higher cloud bills, higher model bills, more rationing, and less predictable economics at the application layer.
Memory is an especially important part of the story because it is not always visible to software buyers until the bill arrives. Reuters reported in January that Samsung and SK Hynix were warning of an acute chip shortage as AI demand soaked up high-value memory supply, leaving other device makers under pressure and pushing the problem into 2026 and potentially beyond. Reuters also reported that Apple was already warning that rising memory costs were starting to bite as suppliers prioritized AI-related chips. In other words, the price of AI is not just a function of model quality or software design. It is also a function of physical component scarcity. When memory is tight, the cost does not stay trapped in the semiconductor supply chain. It moves downstream into cloud pricing, service limits, infrastructure budgets, and token economics. AI may feel intangible at the user interface, but its cost base is increasingly physical.
The same logic now applies to CPUs, networking, and power. Reuters reported ahead of Nvidia’s March conference that as AI moves toward inference and agentic systems, the bottleneck is shifting toward “agent orchestration,” an area where CPUs matter more than the old training-centric narrative suggested. That matters because most enterprises are not deploying one giant model query; they are building layered systems that retrieve data, call tools, pass tasks between agents, manage permissions, and keep humans in the loop. Those workflows consume the whole stack. And when the whole stack matters, the margin promise becomes harder to realize quickly. A coding agent may draft code, but a human still scopes the problem, reviews the output, tests it, and owns the consequences. A support agent may close routine tickets, but exceptions still escalate. So the company often pays for both layers simultaneously: the human layer has not disappeared, and the AI layer has arrived in full. That is why the economic risk is not that AI fails to raise productivity. It is that the cost of deploying AI at scale rises fast enough to delay, dilute, or offset the margin gains management teams are underwriting.
That is the real caution embedded in the Uber example. The lesson is not that AI does not work. It is that AI works well enough to be used constantly, and once it is used constantly, it starts to expose the mismatch between labor-replacement narratives and current cost realities. Unless model efficiency improves materially, unless infrastructure constraints ease, and unless enterprises get much better at governing usage, many companies may discover that they have not yet removed one labor bill so much as added a second one. The old cost arrives in salaries and benefits. The new cost arrives in tokens, memory prices, cloud bills, and megawatts. Over time, that may still prove worthwhile. But in the near term, investors and management teams should be careful with any margin story that assumes AI adoption automatically translates into lower operating expense. With the Wall Street Journal reporting that computing power is already being rationed and Reuters documenting that memory shortages are still feeding through the system, the more realistic conclusion is that AI adoption can remain strategically necessary even while it pressures margins in the interim. Right now, for many enterprises, that is the more honest way to think about the trade-off.