Over the last year, most of my writing has focused on everything outside of NVIDIA. Not because NVIDIA stopped mattering, but because the AI story broadened. Over the past few weeks, one question has kept coming up from readers: when are you going to write about the 800V transition? That question is the right one, because 800V is not a small technical detail. It is one of the clearest signs that AI is moving from a chip-cycle story into an industrial-infrastructure story.
As model capabilities accelerated and Opus 4.5 opened the door to agents, the best and least owned opportunities were beyond the accelerator itself and toward the hardware companies needed to support the agentic explosion: chips, optical, power, cooling, memory, networking, materials and infrastructure. Since October 31, 2025, just before the release of Opus 4.5, NVIDIA is up only about 2% through the close on June 16. In my 100-name Agentic thematic portfolio, it ranks 84th over that period. The index itself is up nearly 70%, and 26 names have more than doubled.
That relative performance matters because it tells us something about the phase of the cycle. The first stage of AI was about GPU scarcity. The next stage broadened into the ecosystem needed to make agents, inference, and applications real. But now, as the industry runs into physical bottlenecks, rising token costs, and a more difficult grid environment, it is time to come back to NVIDIA from a different angle. The next leg is not simply about more GPUs. It is about NVIDIA’s attempt to improve tokens per watt and build the industrial architecture required for AI factories.
Elon Musk’s description of NVIDIA’s Vera Rubin as “a rocket engine for AI” is more than a good sound bite. The full comment captures the moment precisely: “NVIDIA Rubin will be a rocket engine for AI. If you want to train and deploy frontier models at scale, this is the infrastructure you use and Rubin will remind the world that NVIDIA is the gold standard.” That quote signals that the AI buildout has stopped being only a chip-cycle story and has become an industrial-infrastructure story. A rocket engine is not valuable in isolation. It only matters if the entire vehicle can handle the thrust. The fuel system, cooling system, materials, guidance, power flow, and launch structure all have to be designed around the force the engine creates. That is exactly where AI is heading. Vera Rubin may provide the next leap in compute thrust, but that thrust can only be used if the data center around it evolves into something closer to a factory: dense, power-hungry, liquid-cooled, networked, and engineered as one integrated production system.
This paper is not about debating AGI timelines. That conversation matters, but it can too easily drift into abstraction, hype, and speculation. The more immediate reality remains physical. Frontier AI is becoming constrained by power, cooling, networking, memory, land, transformers, substations, and the ability to turn electricity into usable compute. The winner of the next phase will not simply be the company with the cleverest algorithms or the largest model announcement. It will be the company that can build the most efficient physical factory to run those models at scale. That is the brutal hardware bottleneck underneath the AI race.
The key metric underneath this transition is what Jensen Huang has said all year beginning with CES, tokens per watt. In the first phase of AI, the market cared about who had the most GPUs. In the next phase, that question becomes too narrow. The better question is who can produce the most useful intelligence from each unit of electricity. As the cost of the AI buildout rises, as frontier models become more expensive to train and serve, and as the grid becomes harder to access, tokens per watt becomes one of the most important measures of competitive advantage. The organizations that can lower the energy cost of each token can scale faster, price more aggressively, preserve margins, and deploy AI more broadly.
Vera Rubin marks an inflection in this metric. NVIDIA says Vera Rubin NVL72 delivers AI inference at one-tenth the cost per million tokens compared with GB200 NVL72 and trains mixture-of-experts models with one-fourth the number of GPUs compared with GB200 NVL72. NVIDIA also describes the platform as delivering up to 10 times more tokens per megawatt than GB200 NVL72, scaling intelligence within the same power footprint. These are not incremental gains. They represent a step-change in how much useful intelligence can be extracted from a given watt-hour. The Vera CPU is engineered for high-bandwidth, energy-efficient data movement in large AI factories, while the surrounding networking fabric is being redesigned for better efficiency, resiliency, and uptime. The efficiency improvements are real, but they only translate into economic advantage if the surrounding infrastructure can deliver power and remove heat at the densities Rubin enables.
This matters even more as AI moves from the lab into the enterprise. Training gets the headlines, but inference is where AI becomes an operating expense. Every customer query, coding assistant response, agentic workflow, video generation, robotics command, and enterprise automation task consumes compute and power. If AI becomes embedded in every business process, then the cost of producing tokens becomes a daily economic constraint. This has been a headline story now for the last month as costs have risen sharply. For enterprises, the question will not simply be whether a model is powerful enough. It will be whether that model can be run cheaply, reliably, privately, and close enough to the user or workflow to justify the cost.
That is why the edge makes tokens per watt even more important. As companies move more AI workloads closer to factories, hospitals, retail locations, vehicles, robots, phones, PCs, and private enterprise data centers, they will not have the luxury of unlimited hyperscale power. Edge environments are more constrained by space, cooling, energy budgets, and cost discipline. They need useful intelligence in smaller, more efficient packages. The same logic that pushes hyperscale AI factories toward higher-voltage power delivery also pushes enterprise AI toward better power efficiency everywhere in the stack. The future of AI is not only bigger models in larger data centers. It is also more efficient intelligence distributed across the physical economy.
That transition becomes even more critical as AI moves from software into embodied systems. In a hyperscale data center, weak tokens-per-watt efficiency shows up as higher cost, lower margins, more capex, and harder scaling. In embodied AI, it becomes a physical limitation. A humanoid robot, drone, autonomous vehicle, or industrial machine cannot rely on unlimited grid power or hyperscale cooling. It has to perceive, reason, plan, move, balance, manipulate objects, and interact with the world inside a limited energy envelope. That makes tokens per watt not only a measure of economic efficiency, but a measure of autonomy, uptime, and usefulness in the real world.
This is why Vera Rubin should be viewed as the first stage of a much larger power-efficiency theme. Rubin addresses the problem at the source: the industrial-scale production of intelligence inside the AI factory. But as that intelligence moves outward into enterprises, edge devices, robots, vehicles, and physical machines, the same constraint follows it. The AI factory needs tokens per watt to make frontier intelligence economically scalable. Embodied AI needs tokens per watt to make intelligence physically deployable. In the cloud, tokens per watt determines cost. In the physical world, tokens per watt determines whether the machine can work.
This is why the move to 800V DC power delivery is so important. The first phase of the AI boom was easy to explain: the world needed more GPUs. A100 became H100. H100 became Blackwell. Blackwell now gives way to Vera Rubin. But each generation has made the supporting infrastructure more important, not less. More powerful chips create more demand for power, more heat, more memory bandwidth, more networking, and more physical integration. At some point, the bottleneck is no longer simply whether the chip can be produced. The bottleneck becomes whether the customer can deploy the chip at scale inside a purpose-built factory. That is the beginning of the AI factory era.
Vera Rubin marks that transition because the rack is becoming the basic unit of intelligence production. The Vera Rubin NVL72 is not just a faster accelerator inside a normal server. It is a rack-scale AI supercomputer that combines Rubin GPUs, Vera CPUs, sixth-generation NVLink, networking, DPUs, memory, cooling integration, and software into a single production architecture. NVIDIA says NVLink 6 delivers 3.6 terabytes per second of bandwidth per GPU and 260 terabytes per second of scale-up bandwidth per rack. In the old cloud model, a data center was a building filled with servers. In the AI factory model, the rack is the machine, the cluster is the assembly line, and the data center is the plant. The output is not steel, cars, or chemicals. The output is tokens, reasoning, code, agents, video, robotics control, scientific discovery, and automated work.
The modular, cable-free tray design of these rack-scale systems is not cosmetic. NVIDIA says the Vera Rubin NVL72 design can reduce compute-tray assembly time from nearly two hours to about five minutes, with much faster assembly and serviceability versus Blackwell. That matters because when power density rises and every rack becomes a small factory in itself, deployment speed, maintenance, and uptime become core operating advantages. The AI factory is not just about peak performance. It is about keeping expensive compute online, productive, and economically useful.
That is where 800V DC power architecture becomes the hidden foundation of the next AI cycle. If Rubin is the rocket engine, 800V is part of the fuel-delivery system. Lower-voltage architectures force the system to push enormous current through copper, busbars, cables, and power shelves. The more current you move, the more heat you create, the more copper you need, the more space you consume, and the more energy you lose before it reaches the chip. At AI factory scale, those losses are not small. They are lost megawatts, lost rack space, lost uptime, and lost token output. In a world already constrained by electricity, the ability to convert scarce power into useful intelligence becomes a competitive advantage.
The critical point is that higher-voltage DC architectures do not solve the grid problem by creating new electricity. They solve a different but equally important problem: they make each available megawatt more usable. The AI industry is entering a period where power availability will determine who can keep scaling. Utilities, substations, transformers, turbines, transmission lines, permitting, and local politics all matter. Data centers are expected to consume between 6.7% and 12% by 2028 and local pushback on the buildout is growing. But once an organization secures power, the next question becomes how efficiently that power can be delivered to the GPUs and converted into useful output.
A higher-voltage architecture reduces current dramatically. At the power levels now required, moving from 54V rack distribution to 800V DC reduces the current needed to deliver the same power by roughly a factor of fifteen. Because resistive losses scale with the square of current, the theoretical reduction in conductor losses is more than two orders of magnitude before accounting for real-world design tradeoffs such as conductor sizing, protection equipment, conversion stages, and thermal design. NVIDIA has said the physics of using 54V DC in a single 1MW rack could require up to 200 kilograms of copper busbar. At AI factory scale, current is not an abstraction. It becomes copper, heat, weight, routing complexity, rack space, and lost efficiency.
Complete 800V DC reference architectures tailored for next-generation AI platforms are already moving from concept toward ecosystem development. NVIDIA has said it is leading the transition to 800V DC data-center power infrastructure to support 1MW IT racks and beyond starting in 2027, in collaboration with silicon providers, power-system component suppliers, and data-center power-system partners. The significance is not that every facility converts overnight. The significance is that the direction of travel is clear. Power delivery, not just silicon, is becoming the binding constraint on scaling.
Rack power consumption illustrates why the timing matters. Today’s high-end AI racks already sit far above traditional data-center averages, and NVIDIA says current 54V rack power systems begin to strain as racks exceed 200 kilowatts. Future AI factory designs are moving toward several-hundred-kilowatt and eventually megawatt-class racks. Whether the exact system lands at 200 kilowatts, 600 kilowatts, or 1 megawatt, the direction is what matters: power density is rising fast enough that traditional lower-voltage distribution becomes physically and economically strained. The physics of current, copper losses, heat generation, and space force a change in architecture. Higher-voltage DC is not an incremental improvement; it is an enabling condition for the next step in rack-scale density.
Liquid cooling is the other critical gating factor. 800V can deliver the power, but the heat still has to leave the rack. As AI systems move toward higher density, the cooling stack becomes a precision-engineered supply chain: cold plates, manifolds, pumps, heat exchangers, coolant distribution units, leak detection, quick-disconnect valves, and coolant chemistry. This is not traditional air-cooled data-center infrastructure with a minor upgrade. A leak or failure inside a multi-million-dollar rack can become an uptime, safety, and capital-loss event. That makes reliability, qualification, and serviceability part of the investment thesis. The AI factory is not only an electrical architecture. It is also a thermal architecture.
This creates a widening divide inside data-center real estate. Greenfield facilities engineered natively for megawatt-class, liquid-cooled, higher-voltage AI infrastructure should command a premium because they are built around the physical needs of frontier AI factories. Brownfield facilities designed for lower-density cloud workloads may still have value, but they face a harder retrofit question. If they cannot support the power density, cooling loops, floor loading, electrical distribution, and serviceability required by Rubin-class systems, they risk being pushed down the value stack toward less demanding workloads.
This is where 800V connects the macro AI story to the operating reality of the next decade. It does not create new electricity. It improves how efficiently scarce electricity becomes intelligence. By reducing current, copper intensity, heat, conversion loss, and wasted space, higher-voltage architectures help improve the economics of tokens per watt at the factory level. That matters because AI is becoming a cost-curve business. The winners will be the organizations that can produce more intelligence per dollar, per watt, per rack, and per square foot.
The investment map around AI therefore has to broaden, even before naming individual participants. If the next phase is about AI factories, then the relevant supply chain is no longer limited to accelerators. The factory needs electrical infrastructure, high-voltage distribution, power conversion, thermal management, cooling systems, memory, networking, advanced packaging, materials, systems integration, and construction capacity. Vera Rubin pulls the entire industrial stack into the AI story. Every bottleneck around the chip becomes part of the AI supply chain. The market spent the first phase asking who had the accelerators. The next phase asks who can build the factory around them.
The company map around Vera Rubin will overlap with the broader data-center buildout theme, but it will not be identical. Some of the same companies will matter more because Rubin increases the importance of power density, liquid cooling, advanced packaging, and high-voltage distribution. Other names will enter the discussion because the factory is changing. I will map those beneficiaries in a separate piece, but the key point for this paper is that the AI supply chain is moving deeper into the hidden infrastructure layers that make tokens per watt possible.
One of the most overlooked of those hidden layers is chemicals and advanced materials. Higher tokens per watt does not mean the AI factory becomes less materials-intensive. In many ways, it means the opposite. Rubin’s efficiency gains are enabled by more sophisticated materials science across the stack: leading-edge wafer processes, CMP slurries, deposition and etch chemistries, advanced cleaning, HBM, packaging materials, underfills, substrates, thermal interface materials, coolant chemistry, corrosion inhibitors, filtration, and high-voltage insulation materials. The same move that improves the energy cost of each token also raises the importance of the materials that allow chips, packages, racks, and cooling loops to operate at higher density. This is why chemicals belong inside the tokens-per-watt theme. They are not the visible face of AI, but they are part of the hidden layer that makes the efficiency gains physically possible.
For investors, the most attractive opportunities may not be in commoditized construction, but in the harder-to-replicate bottlenecks of the factory floor: high-voltage power conversion, protection and switchgear, liquid-cooling components, precision connectors, thermal management, power semiconductors, and the specialty materials that allow dense racks to operate safely and continuously. The deeper the industry moves into AI factories, the more every hidden subsystem around the accelerator becomes part of the profit pool.
That is the real meaning of Musk’s rocket-engine quote. He is acknowledging that frontier AI is now an infrastructure race. The winners will not simply be the labs with the best model architecture or the organizations with the largest purchase orders. The winners will be the ones that can turn electricity into intelligence at scale. Rubin provides the thrust. Higher-voltage power delivery, liquid cooling, networking, memory, and rack-scale integration allow that thrust to be used. Vera Rubin is where tokens per watt becomes an infrastructure metric; embodied AI is where it becomes an autonomy metric. This is why Vera Rubin is such an important moment: it moves AI from the chip era into the factory era, where the scarce resource is not just compute, but the ability to industrialize compute into reliable, scalable, profitable intelligence production.