The metric that matters in AI compute has changed. For two years it was FLOPs - raw floating-point operations, the brute-force measure of how much math a chip could do. Then it was tokens per second - how fast a model could generate output. Now, with the Vera Rubin platform entering production at partners worldwide, NVIDIA is making a different argument entirely: tokens per watt.

The shift is not subtle. FLOPs measure capability. Tokens per second measure speed. Tokens per watt measures economics. It asks: how much does it cost to run this model, not in chip terms but in power terms? And power is the constraint that is actually breaking the industry.

Vera Rubin NVL72 - the rack-scale configuration, 72 GPUs per rack - is now in production at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius. The supply chain spans 350+ factory sites in 30 countries. NVIDIA calls it the largest, most mature rack-scale supply chain ever assembled. That is marketing language, but it is also fact: no previous silicon platform has entered production with this many manufacturing partners at this scale simultaneously.

The Power Problem

Here is why tokens per watt is the metric that matters now. Data centers are running out of power. Not space, not cooling, not silicon - power. The grid cannot keep up. The Trump EPA is simultaneously considering a rule change that would hand states the power to decide how much public input there is on permits for new power plants - including the gas plants and diesel generators that data centers increasingly rely on. The rollback is being driven, in part, by AI firms that need more data centers and cannot wait for the permitting process.

This is the context in which Vera Rubin arrives. If the constraint is power, then the platform that delivers the most tokens per watt wins. CoreWeave's benchmark - 10x more tokens per megawatt than Blackwell on DeepSeek-R1 - is the number that matters. It does not mean Vera Rubin is 10x faster. It means for the same power budget, you get 10x more inference throughput. Or: for the same inference workload, you need 10x less power.

350+
Factory sites across 30 countries in the Vera Rubin supply chain
10x
More tokens per megawatt than Blackwell (CoreWeave benchmark, DeepSeek-R1)
300+
Global partners backing the platform rollout

The political dimension is inseparable from the technical one. The same week NVIDIA announced Vera Rubin's ramp, the EPA held a public hearing on a rule that would let states decide how much - if any - public input there can be on permits for new sources of air pollution. The proposed rollback comes as data centers face growing pushback from communities over emissions and power consumption. The AI firms want more data centers. The neighbors want clean air. The EPA may be about to side with the firms.

Built in Fort Worth

The supply chain story is its own kind of geopolitics. Wistron, a Taiwanese electronics manufacturer, opened an advanced manufacturing plant in Fort Worth, Texas to produce NVIDIA AI systems. The plant is part of the broader onshoring push - bringing semiconductor and system assembly back to the US in response to export controls, tariff threats, and the strategic vulnerability of depending on Taiwan for advanced chip packaging.

Meanwhile, at the AI Summit in San Francisco, South Korean President Jae Myung Lee appeared with NVIDIA to outline the country's AI future. South Korea is one of the 30 countries in the Vera Rubin supply chain. The geopolitics of AI compute is no longer just about who makes the chips. It is about who assembles the racks, who builds the data centers, who generates the power, and who gets a say in whether any of it can be built at all.

* * *

The Unit Economics Shift

If Vera Rubin delivers on the tokens-per-watt promise, the unit economics of running AI models change. The cost of inference has been falling steadily, but it has been falling because chips got faster. Vera Rubin introduces a second axis: the cost of power. If you can get 10x more tokens per megawatt, your marginal cost per token drops not because the chip is cheaper but because the electricity bill is lower.

This matters for the business model of AI itself. The race to build frontier models - Anthropic's Opus 5, OpenAI's GPT-5.6, Google's Gemini updates - is a race that consumes enormous amounts of inference compute. Every query costs power. Every agent running in production costs power. If the power cost per token drops by an order of magnitude, the economics of agentic AI - models that run continuously, making dozens of API calls per task - become viable at scale.

The next platform war in AI will not be won by the model with the best benchmark. It will be won by the platform that delivers the most intelligence per watt. Vera Rubin is NVIDIA's opening bid.