DC Decoded
One Race, Four Winners
The AI infrastructure buildout is not a single market. It is a chain of value pools, and different companies win at different stages.
How I got here
I write DC Decoded, a series on data centers, power infrastructure, and the AI supply chains most people never see.
The deeper I go into this world the more one question keeps surfacing. Everyone is spending. Everyone is building. The announced numbers are staggering, $500 billion here, $200 billion there, governments announcing AI investment in the same breath as national security. But when you step back from the press releases and look at the actual moves, the chip deals, the infrastructure commitments, the partnership structures, something interesting emerges.
This is not a race with one finishing line.
It is four distinct theories of how to win. Each being pursued simultaneously. Each betting the other three need them. And each resting on assumptions about the industry that have not yet been fully tested.
I did not arrive at this from reading analyst reports or following headlines. I arrived at it through DC Decoded, tracing supply chains, mapping infrastructure decisions, understanding what data centers actually cost and how long the hardware inside them actually lasts. The picture that emerged was not the one most coverage suggests.
The four theories
Theory A - Own the full stack OpenAI
OpenAI is attempting something no technology company has ever tried at this scale. Full vertical integration from energy to chip to model to consumer product.
Stargate, the $500 billion infrastructure program built with SoftBank, Oracle, and Microsoft, is not just a data center project. The Abilene Texas facility alone is approaching 5 gigawatts of power consumption. OpenAI is building dedicated energy supply behind the meter, natural gas, solar, and small modular nuclear reactors, because it cannot wait for the grid. SoftBank is converting a shuttered EV factory in Ohio into a modular data center fabrication plant, manufacturing the buildings that house the computers. OpenAI’s custom chip codenamed Titan, co-designed with Broadcom and fabricated on TSMC’s 3nm process, targets mass production in the second half of 2026. Stargate has locked up supply deals with Samsung and SK Hynix for up to 900,000 DRAM wafers per month, roughly 40% of projected global DRAM output.
Energy. Manufacturing. Buildings. Memory. Chips. Data centers. Models. Consumer product. Every layer.
The strategic logic is clear. If you own the infrastructure you cannot be held hostage by any single supplier. No Nvidia pricing power. No Microsoft exclusivity. No Google TPU lock in. You control the cost structure of the entire operation.
The financial reality is harder. OpenAI projects operating losses ballooning to roughly three quarters of revenue by 2028. It expects to burn through roughly 14 times as much cash as Anthropic before turning a profit. Analysts have flagged the company could run out of cash by mid-2027 without continued external funding. Sam Altman has defended the spending plainly. The risk of not having enough compute is more significant than the risk of having too much.
What has to be true for this to work. AI demand grows fast enough and consistently enough to justify the infrastructure before the cash runs out. The custom chip program delivers on schedule. And the vertical integration creates cost and capability advantages large enough to justify the losses incurred building it.
What OpenAI fears most. A more efficient model trained at a fraction of the cost makes the infrastructure advantage irrelevant. DeepSeek proved in early 2025 that training costs can be dramatically reduced through algorithmic efficiency. If that trend continues the $500 billion bet on infrastructure scale may be optimising for the wrong variable.
Theory B - Own the model Anthropic
Anthropic is making the opposite bet. Do not try to own the stack. Be the best model running on infrastructure built by others who compete for your business.
The strategy is deliberate and increasingly well capitalised. Anthropic committed to 1 million Google TPUv7 chips coming online in 2026, roughly 1 gigawatt of compute, in a deal worth approximately $52 billion. Amazon’s Project Rainier runs 500,000 Trainium2 chips for Anthropic scaling toward 1 million. A $50 billion partnership with Fluidstack builds custom data centers in Texas and New York. A $30 billion Series G at a $380 billion valuation provides the capital to sustain it.
The key insight. Anthropic is getting access to Google’s and Amazon’s custom chips without paying the capital cost of building them. Google built TPUs over a decade. Amazon built Trainium over years. Anthropic gets the benefit of both at a cost per token roughly 50% lower than equivalent Nvidia GPU configurations, without the years and billions required to build either.
The financial trajectory reflects this discipline. Anthropic projects dropping cash burn to roughly one third of revenue in 2026 and down to 9% by 2027. It expects to break even by 2028, the same year OpenAI projects its worst losses. Enterprise large language model spend share has surged to 40% in late 2025, up from 24% in 2024 and 12% in 2023. OpenAI’s share has fallen to 27% from 50% in 2023.
What has to be true for this to work. Model quality remains the primary purchasing decision. Enterprises choose the best model and run it wherever the infrastructure is cheapest. The model moat holds even as infrastructure becomes commoditised.
What Anthropic fears most. A competitor builds a better model on cheaper infrastructure and the cost per token advantage gets competed away. Or one of its infrastructure partners decides the model business is too valuable to leave to someone else.
Theory C - Own the platform Google, Amazon, Microsoft, Meta
The hyperscalers are playing the most complex game. They are simultaneously Nvidia’s biggest customers and building Nvidia’s biggest competitors. They are hosting the model companies that may eventually challenge them. And they are investing in custom chip programs whose success determines whether their current infrastructure investments hold their value.
The strategic logic. Own the infrastructure layer that nobody can skip. Build custom chips to reduce Nvidia dependency. Rent compute to whoever builds the best models. Be the cloud that AI runs on whether OpenAI wins or Anthropic wins or someone we have not heard of yet wins.
Google is furthest along. TPUs have been running since 2015. Gemini is trained on TPUs. Google is now selling TPU access to Meta, turning a decade of internal investment into an external revenue stream. The cost of training and inference per token on TPUs is estimated at 30 to 50% lower than competitors using Nvidia GPUs.
Amazon’s Trainium program is the most credible Nvidia alternative for training at scale. Project Rainier, 500,000 Trainium2 chips, is already running Anthropic’s models. Trainium3 is confirmed for both training and inference workloads by Anthropic and OpenAI. When the companies building the most important models choose your custom chip over Nvidia for production workloads that is a credibility signal no marketing can replicate. Amazon’s custom silicon business has grown into a $10 billion plus annual run rate within AWS.
Microsoft is the most conflicted. Deepest Nvidia dependency. Biggest OpenAI investment. Its Maia chip program targeting inference is deployed in Azure data centers but behind the ambition. Most exposed if the assumptions underlying current Nvidia hardware investments prove optimistic.
Meta is moving fastest on custom silicon. Four generations of MTIA chips announced in March 2026 built on RISC-V with 25 times compute gains across the lineup. A new chip generation every six months. Simultaneously buying Google TPUs. Diversifying faster than any headline suggests.
What has to be true for this to work. The custom chip programs mature fast enough to reduce Nvidia dependency meaningfully. Model companies keep running on hyperscaler infrastructure rather than building their own at scale. And the platform position proves more durable than the model position.
What the hyperscalers fear most. OpenAI’s Stargate succeeds and creates a vertically integrated competitor that no longer needs their cloud. Or the inference market migrates to purpose built chips faster than the current infrastructure investments can adapt.
Theory D - Own the chip Nvidia
Nvidia did not plan to dominate AI. A gaming GPU built for parallel processing turned out to be exactly what neural network training required. That accident gave Jensen Huang a ten year head start and he moved fast to lock it in.
The CUDA software ecosystem is fifteen years old. Every AI researcher learns it. Every major framework runs on it. Switching to a competitor means rewriting code, retraining teams, accepting performance uncertainty. That switching cost is Nvidia’s deepest moat, deeper than the hardware itself.
Annual chip releases keep the performance gap wide enough that the calculation always favours staying. Hopper to Blackwell to Rubin, each generation two to three times more powerful than the last. By the time a competitor’s custom chip is ready to challenge the current generation Nvidia has already shipped the next one.
Nvidia also moved on inference, the area where its dominance was most threatened. Blackwell Ultra delivers 15 petaFLOPS of inference compute. Rubin arriving in the second half of 2026 promises 50 petaFLOPS, five times Blackwell. Nvidia launched Dynamo, an open source inference framework boosting throughput 30 times on Blackwell hardware. And a $20 billion licensing deal with Groq absorbed a potential inference disruptor into Nvidia’s ecosystem rather than competing with it.
Revenue reflects the position. Fiscal year 2026 hit $215.9 billion, up 65% year on year. Data center revenue at $193.7 billion. Operating margins above 70%.
What has to be true for this to work. CUDA’s switching costs hold as open source inference frameworks mature. Training demand does not normalise so fast that the inference migration happens before Nvidia has secured that market. And the custom chip programs at the hyperscalers take long enough to develop that Nvidia has time to compete on inference before they arrive at scale.
What Nvidia fears most. Inference at scale migrates to custom ASICs faster than training revenue sustains the valuation. Custom ASIC shipments are projected to grow at 44.6% in 2026 versus GPU shipments at 16.1%. If inference goes to ASICs Nvidia’s addressable market shrinks, not because it lost but because the market it dominated becomes a smaller share of the total.
The inference shift where all four theories get tested
Training a large model happens once. Or a handful of times. It is expensive and it favours the biggest GPU clusters.
Inference, serving that model to millions of users every second of every day, never stops. And inference has completely different hardware requirements. Low latency. High throughput. Energy efficiency. Cost per token.
Inference now accounts for roughly two thirds of all AI compute, up from one third in 2023. The inference market is projected to exceed $50 billion in 2026 growing faster than training compute for the first time. Purpose built inference chips deliver three to five times better cost per token than training optimised Nvidia GPUs for serving workloads.
This shift is where all four theories get stress tested simultaneously.
If inference migrates to custom ASICs, Google TPUs, Amazon Inferentia, Groq LPUs, Nvidia’s addressable market narrows. The hyperscaler custom chip programs become more valuable faster than the depreciation schedules on current Nvidia hardware assumed. Anthropic’s strategy of running on the cheapest available inference infrastructure compounds in its favour. And OpenAI’s vertical integration bet either proves its value, owning the inference infrastructure at scale, or becomes its vulnerability if the Titan chip arrives too late.
The inference shift is not a threat to AI infrastructure spending. It is a restructuring of who captures the value from that spending.
Why this does not have to be winner takes all
Most coverage frames AI as a race with one finishing line. One company will win. The others will be footnotes.
I think that framing is wrong.
The internet did not produce one winner. It produced winners at each layer of the stack. Each layer had different economics, different switching costs, different competitive dynamics. The layered structure created durable positions at every level and made the whole system more resilient than any single player could have built alone.
AI infrastructure is developing the same layered structure. A chip layer. An infrastructure layer. A model layer. A consumer and agentic layer. If Nvidia owns the training chip layer, hyperscalers own the infrastructure platform layer, Anthropic owns the enterprise model layer, and OpenAI owns the consumer and agentic layer, all four can exist simultaneously. All four need each other. All four compete fiercely within their own layer while depending on the others to function.
That is probably the best outcome. For the industry. For innovation. For the enterprises building on top of all of it. And for the people who use these products without knowing or caring which chip their query ran on.
What could break it
The four layer structure is not inevitable. It requires each player to stay within the logic of their own bet.
OpenAI is the most likely to break the structure intentionally. The entire Stargate thesis is that owning every layer is the winning position. If it succeeds it validates the full stack theory and puts pressure on every other layer simultaneously.
The inference shift is the most likely to break it unintentionally. If inference migrates to custom ASICs faster than current Nvidia hardware assumptions hold the economics underpinning the hyperscaler position become fragile in ways not yet visible in reported earnings.
And the model layer is the most contested, because unlike chips or data centers model quality is not protected by capital intensity. A small team with the right insight and the right compute access can build a model that displaces a larger competitor faster than any physical infrastructure advantage can respond.
The race is real. The spending is real. The stakes are genuinely historic.
But the finishing line may not be a single winner.
It may be four companies, each owning one layer of the most consequential infrastructure ever built, each making the others possible.
This series started with a podcast featuring Edward Fishman, whose thinking on invisible dependencies and economic chokepoints sent me down this rabbit hole. If you have not read Part One, on helium, supply chains, and what the Strait of Hormuz revealed that the oil conversation missed, it is linked here.
Sources: OpenAI SEC filings and press releases, Anthropic CNBC reporting, Deloitte TMT Predictions 2026, TrendForce 2026, McKinsey Global Institute, Goldman Sachs Research, Company earnings calls from Amazon, Google, Microsoft, Meta, Nvidia, Reuters, Bloomberg, CNBC, Fortune, Wall Street Journal, 2025 and 2026.
*Analysis based on publicly available information as of March 2026. This is a rapidly evolving space and positions may have shifted since publication.
#DCDecoded #AIInfrastructure #Semiconductors #Strategy #Nvidia #OpenAI #Anthropic #Leadership #AIChips