DC Decoded
Training vs. Inference: The Compute Shift That Matters
Why the centre of gravity in AI infrastructure may move from building models to running them at scale.
By Praveen Gangaraju · Companion to LinkedIn Post #3 · All figures based on publicly available data as of March 2026
When I started digging into why data center power demand keeps rising, even as AI models are genuinely getting more efficient, I expected a simple answer. Bigger models, more training compute, more power. It made sense. The data told a different story.
WHAT TRAINING IS AND WHAT IT COSTS
Training is when an AI model is built from scratch. Compute-intensive and expensive.
— GPT-3 (2020): ~$4.6M (OpenAI / Lambda AI estimates)
— GPT-4 (2023): ~$100M+ (industry estimates, widely cited)
— GPT-5 (2025): used less training compute than GPT-4 (Epoch AI)
The scaling-equals-bigger pattern broke. But training happens once. The run ends. Training is a sprint.

WHAT INFERENCE IS AND WHY IT CHANGES EVERYTHING
Inference is every time anyone uses AI. Every question. Every response. Without ever stopping.
A single GPT-4o query uses approximately 0.34 Wh (OpenAI, Epoch AI Feb 2025). Less than a lightbulb in a minute.
ChatGPT: 900 million weekly active users as of February 27, 2026 (OpenAI confirmed). 2.5 billion prompts every day. The cumulative inference energy overtakes training GPT-4 within months. Then it compounds permanently.
80–90% of all AI computing power today is consumed by inference, not training (MIT Technology Review / Lawrence Berkeley National Lab).

THE EFFICIENCY PARADOX
If models are getting more efficient, why does power demand keep rising?
Google: 33x less energy per Gemini query in one year. GPT-5 used less training compute than GPT-4. Efficiency is real.
But adoption is outrunning it. ChatGPT: 400M to 900M weekly users in twelve months.

THE REASONING SHIFT
GPT-5 averages ~18.9 Wh per medium-length prompt vs ~0.34 Wh for simpler earlier-model queries (University of Rhode Island AI Lab, 2025). Range: 2–45 Wh. Early research, direction is clear. As reasoning becomes default, compute cost per user rises even as training efficiency improves.
Three forces simultaneously: efficiency improving, adoption outpacing it, reasoning multiplying cost per interaction.
WHAT THIS MEANS
The industry is building for the 900 million people using AI today and billions more tomorrow, for harder tasks, every day. Not for training runs.
I spent years in commercial conversations in Singapore, Malaysia, and India. Operators were never asking about training. They were asking about what their users would need tomorrow. That instinct was right.
The industry is not building for AI models. It is building for AI users. And that race has no visible finish line.
--—————————————————————————————————————-
All figures based on publicly available data as of March 2026. This is a fast-moving field, numbers will evolve. Where estimates vary, I’ve used the most widely cited figures and flagged uncertainty. Nothing here is precise engineering data, it is an honest attempt to understand a system the industry itself does not fully disclose. This content is original analysis, all source data is cited and publicly available.