Agentic AI and the Broadening of Infrastructure Economics

Executive Summary
As AI evolves from answering single prompts to executing multi-step, tool-using workflows, the marginal unit of infrastructure demand is fundamentally transforming. The true implication of the “whole rack” thesis is not that GPUs are losing their relevance, but rather that they now require a balanced, integrated system to function at scale. In this agentic world, CPUs regain relevance as orchestration engines, networking becomes more valuable because agentic systems create more east-west traffic, power and cooling move from background line items into binding deployment constraints, and enterprise server OEMs benefit because customers need validated, AI-ready rack-scale systems rather than loose component purchases. The five core thematic areas that matter most are: power/cooling, networking/interconnect, CPU orchestration, enterprise server refresh and rack integration, and HBM plus advanced packaging. These are the lanes where economics broaden first if agentic deployments move from pilots to production over the next 4–8 quarters.
The investment implication is more selective than the bullish narrative suggests. Value is still likely to concentrate where scarcity power, qualification barriers, and switching costs are highest. The main risk is not a macro shock alone; it is that enterprise agentic adoption remains stuck in pilot mode, in which case the thesis narrows back toward a more GPU-centric story and the non-GPU broadening beneficiaries lose momentum.
Introduction
The core mechanism driving the “whole rack” thesis is a structural change in the composition of demand, not just a broad inflation of infrastructure spending. While a traditional chatbot can operate efficiently on an accelerator-heavy stack with limited persistence, modest orchestration, and relatively simple data flows, an agentic system breaks that mold. Because it must plan, invoke tools, access enterprise files, query databases, verify outputs, maintain state, and loop through multiple steps to deliver a usable result, this operational complexity forces the economic relevance outward from the GPU to the rest of the rack: CPUs, memory, storage, networking, power, cooling, and systems integration.
That framing matters because since the launch of ChatGPT, the market has often treated AI infrastructure as a single-bottleneck story centered on accelerators. This belief is too narrow for the next phase. They point to a causal chain in which GPU inference demand rises first, then CPU orchestration becomes more important, then storage and networking requirements expand, and finally facility constraints such as power delivery and cooling emerge as the hard physical ceiling on deployment. In other words, the thesis only works if AI becomes operational, not just computational.
There are several reasons the timing matters now, largely confirmed by recent commercial data. Commentary exiting 2025 and entering 2026 from major ecosystem players—including AMD, Intel, Nvidia, and Dell—signals a definitive inflection point: the growth of inference is now exceeding training, and traditional compute demand is surging specifically to handle AI orchestration. This urgency is further amplified by Jensen Huang’s recent remarks on token multiplication and the rapid rollout of open-source agentic frameworks like OpenClaw. First, hyperscaler capex has already reached a scale where small shifts in rack composition translate into very large changes in revenue pools across semiconductors, networking, and facilities. Second, the enterprise installed base is old enough that AI readiness can piggyback on a replacement cycle rather than requiring entirely new greenfield budgets. Third, rack power density has risen fast enough that power and cooling are no longer background engineering issues; they are deployment gates. And fourth, the reports consistently argue that inference and orchestration are becoming more durable demand drivers than a one-time training cycle alone.
Still, the expanded “whole rack” thesis needs some skepticism. The bear case is not that AI infrastructure collapses outright. It is that the broadening fails to materialize at the scale or speed implied. If agentic projects are cancelled, delayed, or remain narrow; if GPU software optimization absorbs more of the workload than expected; if hyperscalers internalize more of the profit pool through custom silicon and integrated stacks; or if enterprise refresh demand fades before agentic demand replaces it, then much of the non-GPU enthusiasm proves premature. That is why the best way to use this framework is not as a blanket endorsement of “AI adjacencies,” but as a ranked map of which layers are most likely to capture value.
For capital allocators, a concise framework for evaluating these five verticals is as follows:
- Power, electrical infrastructure, and cooling
- Why it is core: Physical deployment bottleneck; longest lead times; high qualification barriers.
- Networking and interconnect
- Why it is core: Agentic workflows increase east-west traffic and switching intensity.
- CPU orchestration and heterogeneous compute
- Why it is core: Tool use, data prep, state handling, and control-plane work increase CPU relevance.
- Enterprise server refresh and rack integration
- Why it is core: Aging installed base plus need for validated AI-ready systems.
- HBM, memory, and advanced packaging
- Why it is core: Upstream choke points still govern system availability and economics.
1. Power, Electrical Infrastructure, and Liquid Cooling
If one vertical deserves to move from “important” to “core,” it is power and thermal management. The reports are unusually consistent on this point. They describe power and cooling not as derivative beneficiaries of AI buildout, but as the layer that can physically cap deployment regardless of how strong silicon demand appears on paper. As rack densities rise and more workloads run continuously rather than episodically, average server utilization, power draw, and thermal stress all increase. In that setup, the scarce asset is not simply compute; it is the ability to turn compute into deployed, operational capacity.
This is why power and cooling matter more than generic “data center exposure.” Scarcity sits in a few specific places: switchgear, UPS systems, liquid cooling distribution units, transformers, interconnection access, and qualified thermal architectures that hyperscalers are willing to standardize around. These are engineering-heavy categories with real lead times and nontrivial qualification cycles. That combination tends to create better pricing power than a simple volume expansion story in servers or commodity components. The structural reason this vertical matters is that it is harder to disintermediate than many other parts of the stack. A hyperscaler can build a custom ASIC, and a large cloud provider can internalize some server design or network architecture. It is much harder to wish away power conversion, heat rejection, transformer lead times, or liquid-cooling implementation at high densities. Timing is the major uncertainty. Physical infrastructure can be the correct thesis but still disappoint on cadence if utility bottlenecks, permitting, or project sequencing push revenue recognition further out than the market expects.
The real insight here is that agentic AI turns racks from bursty, training-style loads into always-on digital labor. A chatbot inference cycle might spike for seconds and then idle; an agent, however, runs continuous loops, planning, tool-calling, verifying, and handing off work, 24/7 inside enterprise workflows. That sustained utilization pushes average rack power draw dramatically higher and forces a wholesale shift toward liquid cooling and denser electrical infrastructure simply to keep the agents from thermally throttling or tripping facility limits. The transition is no longer optional; it becomes the gating factor that determines whether thousands of agents can actually be deployed at production scale.
2. Networking and East-West Interconnect
Networking is the most compelling non-GPU broadening theme because the mechanism is intuitive and the market structure is attractive. Agentic systems generate more data motion inside the data center than simple prompt-response inference. The model does not just produce an answer; it queries data, calls tools, exchanges context, routes intermediate steps, and often coordinates across multiple processes. That makes low-latency switching, scale-out fabrics, and interconnect quality more economically important. The reports repeatedly describe this as a shift toward higher east-west traffic intensity, which is precisely where the best networking names are positioned.
This area is also attractive because it combines demand broadening with a relatively concentrated value chain. One report explicitly describes networking systems as tending toward a winner-take-most structure while network silicon looks more like a duopoly. Cisco has quietly repositioned itself around the real bottleneck of agentic AI: security, control, and orchestration of autonomous workflows, rather than just bandwidth. This shift was initially underappreciated, especially as Gartner’s 2025 Security Operations Hype Cycle placed AI assistants and AI SOC agents at the “Peak of Inflated Expectations.” At the time, the dominant view was that enterprise experimentation with AI agents lacked clear ROI and introduced new attack surfaces, with more than half of successful attacks expected to exploit weaknesses in access control and prompt-layer vulnerabilities.
The deeper insight is that every agentic action creates a cascade of internal network hops that a chatbot never needed. Instead of one clean request-response pair, an agent might fetch a document from storage, query a database, invoke three different APIs, cross-check results with another model, and then route the output, all in microseconds and all inside the same rack or across adjacent racks. That multi-hop, east-west pattern multiplies traffic volume and latency sensitivity far beyond what GPU-only inference required, turning the network fabric into the connective tissue that either enables or throttles the entire agentic workflow at scale.
3. CPU Orchestration and Heterogeneous Compute Balance
The CPU branch of the thesis is where the market may still be underappreciating nuance. Intel CFO David Zinsner recently emphasized this point by saying “the CPU has become cool again this year.” The bullish case is not that CPUs retake the center of the AI stack. It is that agentic systems create an orchestration layer around accelerators: tool invocation, retrieval, state handling, data preprocessing, policy checks, memory coordination, and general control-plane work. Several of the uploaded reports explicitly describe CPUs as unexpectedly relevant again, arguing that the orchestration burden rises with multi-step workflows and that CPU supply has become meaningfully tighter than the market expected.
That makes CPUs one of the more interesting broadening beneficiaries, but only if framed properly. The near-term opportunity is strongest where enterprise compatibility and installed-base inertia matter, which favors x86 vendors. The medium-term threat is that hyperscalers continue shifting toward Arm and custom silicon, which internalizes some of the highest-value CPU demand and weakens the idea of a broad, durable merchant CPU renaissance. This is also one of the areas where expectations need to be kept under control. The safer interpretation is that even a more modest increase in workflow complexity can make CPU attach and orchestration economically meaningful. You do not need the most extreme demand assumptions for this part of the thesis to work. But you do need evidence that enterprises are deploying real workloads with material control-plane and coordination overhead rather than just running more accelerator-heavy inference.
The core insight is that the agent’s intelligence lives in the orchestration layer, not just the model. While the GPU still handles the heavy token generation, the CPU must plan the next step, maintain session state across tool calls, enforce enterprise guardrails, prepare context for retrieval, and decide when to hand off to another agent. As workflows move from one-shot prompts to 1,000× token loops with branching logic and persistent memory, this control-plane load grows independently of raw inference volume, turning CPUs from background support into a co-equal requirement that keeps the entire rack balanced and productive.
4. Enterprise Server Refresh and Rack Integration
Within this vertical, the broad thesis is that enterprise AI deployment will often travel through a refresh cycle rather than through entirely new stand-alone AI budgets. That matters because it brings the theme out of hyperscaler capex and into mainstream IT spending. The reports repeatedly argue that the installed base is old, that newer server generations offer attractive consolidation economics, and that the market may be underestimating how much AI readiness can attach to a normal replacement decision.
While this vertical sometimes faces margin pressure on the commoditized server side, the critical differentiator in the agentic era is high-performance enterprise storage. As an example, Pure Storage represents a crucial structural beneficiary here. As agentic AI moves from pilots to production, these workflows rely continuously on Retrieval-Augmented Generation (RAG). RAG effectively makes a low-latency data platform as critical as compute; thousands of autonomous agents must constantly retrieve files, query vector databases, and integrate unstructured corporate data to perform useful work. Any storage latency at the systems layer instantly bottlenecks the expensive GPUs, making flash-centric data platforms essential revenue transmission mechanisms that capture higher value than generic integration. In practical terms, this entire vertical proves whether the whole-rack thesis is entering broad enterprise deployment or remaining concentrated in hyperscale and sovereign AI.
The pivotal insight is that legacy servers and spinning-disk storage architectures were never designed for thousands of always-on agents running heterogeneous, data-intensive workloads inside the same rack. Agentic systems demand validated, pre-integrated stacks that combine GPU inference, CPU orchestration, high-speed internal networking, and persistent storage access, none of which bolt cleanly onto five-year-old hardware. Furthermore, as agents transition into production, enterprise security mandates and the inherent risk of model hallucination or misuse are driving a hard requirement for the physical control of these workflows. While agents may query the cloud, the core orchestration layer, the integration with private enterprise data, and the application of security guardrails must reside on-premises or in a highly controlled hybrid environment. The refresh cycle therefore becomes the practical on-ramp: enterprises replace outdated servers and storage not just for cost savings but to unlock a secure, physically managed data platform capable of hosting digital labor at scale, turning routine capex into the structural bridge that carries agentic AI from pilot to production.
5. HBM, Memory, and Advanced Packaging
The final core thematic area is upstream and the one that gets most of the press these days, memory. Memory matters because the whole-rack thesis still depends on the stack being physically buildable. In that sense, HBM and advanced packaging remain central even if they are more consensus than the non-GPU broadening story. The reports consistently identify memory and packaging as among the tightest choke points in the ecosystem, with strong value capture because technical complexity limits entrants and because these layers gate delivery for the rest of the rack.
This vertical matters for two reasons. First, it still controls availability. A beautifully specified AI rack does not ship at scale if HBM remains constrained or if advanced packaging capacity is too tight. Second, it still controls a large part of the profit pool. Even if the market wants to move on from “just own the obvious bottlenecks,” the reality is that HBM and advanced packaging continue to command scarcity pricing and strategic importance. The right interpretation is not that these names are no longer attractive; it is that they are more recognized and therefore less differentiated as a variant view than power, networking, or selective OEM refresh.
The real insight is that agentic workflows do not just consume more tokens—they consume more context, faster. Every tool call, every file retrieval, every vector search, and every state hand-off requires ultra-high-bandwidth memory to keep the GPU and CPU layers fed without stalling. As I detailed in my recent analysis, The Molecular Moat: How Chemistry Will Power the Next Decade of AI Infrastructure, this physical buildability increasingly rests on a foundation of specialty polymers. HBM stacking requires advanced non-conductive films, while CoWoS-style advanced packaging depends on organic redistribution layers and epoxy molding compounds to protect the thousands of tiny interconnects inside the stack. Advanced packaging is what glues those heterogeneous components together at rack scale; without sufficient HBM and CoWoS-style capacity the entire system stays supply-constrained. The broadening thesis therefore still rests on this upstream foundation: the rack can only expand into balanced, production-grade agentic capacity if memory and packaging keep pace with the shift from episodic to continuous digital labor.
Pulling the Vertical View Together
The best way to synthesize the expanded paper is to separate broadening of demand from quality of value capture. The whole-rack thesis says demand broadens beyond accelerators. That part of the argument is credible if agentic AI scales. But value does not accrue evenly. The strongest economic logic still sits in bottlenecks and concentrated layers. That is why power/cooling, networking, and upstream memory/packaging often look better than generic integration. CPU orchestration is real but more contested because Arm and custom silicon can absorb some of the upside. OEM refresh is real but lower quality because revenue can rise without a commensurate jump in structural margins.
That is also why the five verticals should not be treated symmetrically. If ranking them by combination of durability, clarity of mechanism, and scarcity power, power/cooling comes first, networking second, CPU orchestration third, enterprise OEM refresh fourth, and HBM/packaging as economically central but less differentiated as a research angle because the market already understands much of that bottleneck. That ranking is broadly consistent with the consolidated synthesis and with the strongest recurring patterns across the uploaded reports.
Conclusion
The expanded version of the paper leaves you with a narrower but stronger conclusion than the original broad narrative. The theme is not “AI infrastructure gets bigger.” The theme is that agentic AI changes what matters inside the rack. The key message is timing as OpenClaw has opened the door to the agentic world. As workloads become more stateful, tool-using, and continuously operating, demand broadens into system balance: CPU, memory, data motion, networking, power delivery, cooling, and validated rack integration. That is the real insight. GPUs remain central, but the marginal economics are no longer exclusively theirs if the workload model shifts the way the reports suggest.
The five core areas that deserve primary attention are power/cooling, networking, CPU orchestration, enterprise server refresh, and HBM plus advanced packaging. The highest-quality broadening exposures are generally those with qualification barriers, scarcity power, or concentrated market structure. The weakest parts of the thesis remain pilot-to-production conversion, the possibility of hyperscaler internalization, and the risk that some of the OEM uplift reflects cycle timing rather than durable structural change.
The associated stock list for this thematic framework is available through your 22V salesperson, or on the 22V website for subscribers.