Here is a strange possibility.
AI could become dramatically more efficient—and data centers could still use more electricity.
At first, that sounds wrong.
If a model needs less computing power for each answer, shouldn’t we need fewer GPUs, smaller server farms, and less electricity?
Sometimes, yes.
But there is another force working in the opposite direction.
When AI becomes cheaper, faster, and small enough to put into more products, we may simply use much more of it.
The question is not only how much energy one AI task uses. It is how many AI tasks the world decides to run.
That distinction could become one of the most important ideas for understanding the next stage of the AI data-center boom.
First, Smaller Does Not Mean Weak
For a while, the AI story sounded simple: bigger model, more computing, better capability.
That logic still matters at the frontier. The most capable models often need enormous training clusters and large inference systems.
But everyday AI does not always need the biggest model available.
A customer-service bot does not need to solve every scientific problem. A phone that rewrites a short message does not need the same system used to train a frontier model. A factory camera may need to recognize a small number of very specific conditions over and over again.
That is where smaller, specialized models become interesting.
Microsoft’s Phi family is built around this idea. Microsoft describes Phi as a family of small language models designed to deliver useful AI with fewer computing resources, including models that can run on PCs and edge devices.
Google is pushing in the same direction. In 2025, it introduced Gemma 3 270M, a 270-million-parameter model designed for efficient, task-specific applications. Google reported that an INT4-quantized version used only 0.75% of a Pixel 9 Pro battery during 25 test conversations.
That does not mean a 270-million-parameter model can replace every large cloud model.
It means we are learning to ask a better question:
What is the smallest model that can do this particular job well enough?
Why Model Efficiency Matters to a Data Center
Imagine that an AI service uses 10 units of computing for one request.
Now engineers improve the model, software, and hardware so that the same useful result needs only 1 unit.
If the number of requests stays the same, the effect is obvious:
10× better efficiency → roughly 90% less computing for that workload
That can mean fewer accelerator-hours, less electricity, less cooling, and lower cost.
This is why optimization matters so much.
Engineers can reduce the cost of inference in many ways: use a smaller model, compress numerical precision, specialize a model for one task, improve the software stack, or run it on hardware designed specifically for AI.
The exact technique is less important here than the direction.
Useful AI per watt is improving.
And the Improvement Is Happening Very Fast
The International Energy Agency’s 2026 Key Questions on Energy and AI report gives us an unusually clear way to see the change.
According to the IEA, the electricity used per individual AI task has been falling extremely quickly. In recent years, software and hardware improvements have reduced energy use per task by at least an order of magnitude annually.
That is a huge improvement.
If efficiency were the only thing changing, we might expect the AI electricity problem to start shrinking.
But that is not what the IEA sees.
In 2025, total data-center electricity consumption grew by about 17%.
Electricity use by AI-focused data centers grew even faster—about 50%.
So we have two facts that look contradictory:
Energy per AI task ↓ sharply
Total AI data-center electricity ↑ sharply
Both can be true at the same time.
The Missing Variable Is Usage
Suppose an AI request becomes ten times cheaper to run.
That might save a company money.
But it can also make completely new products affordable.
A company that used AI once per customer interaction might use it ten times.
A software tool might add AI to every document, every email, every image, and every meeting.
An AI assistant might stop waiting for a human question and begin working continuously as an agent—searching, comparing, planning, checking, and calling other tools.
Suddenly, the unit of demand is no longer “one chatbot answer.”
It might be hundreds of model calls behind one user action.
The IEA highlights exactly this problem. Simple text queries are becoming much more efficient, but new uses such as reasoning, video generation, and agentic tasks can consume hundreds or even thousands of times more energy per query than simple text generation.
So the future depends on three things moving at once:
- Efficiency: how much compute one useful task needs.
- Adoption: how many people and companies use AI.
- Task intensity: how much work each AI application asks the model to do.
The first pushes energy demand down.
The other two can push it up.
This Is Where the Jevons Paradox Enters the Story
There is an old economic idea called the Jevons paradox.
The simple version goes like this:
When a technology becomes more efficient, it becomes cheaper to use.
When it becomes cheaper, people may use much more of it.
Sometimes the increase in usage can offset part—or even all—of the efficiency savings.
That does not happen automatically. It depends on how strongly demand responds to lower cost.
But AI is a good place to watch for this rebound effect because many possible uses have not even been invented yet.
A Simple Example
Imagine a company runs 1 million AI requests per day.
Each request needs 10 units of compute.
Total demand is:
1 million × 10 = 10 million compute units
Now the model becomes ten times more efficient.
Each request needs only 1 unit.
If usage stays at 1 million requests:
1 million × 1 = 1 million compute units
Great. Demand falls by 90%.
But what if cheaper AI lets the company build new features and usage rises to 15 million requests?
Then:
15 million × 1 = 15 million compute units
Each request became ten times cheaper.
Total compute still increased.
This is not a prediction that AI demand must rise this way.
It is simply the calculation we need to keep in mind.
Total compute = compute per task × number of tasks
Small Models Can Also Move AI Out of the Data Center
There is another twist.
Smaller models do not only make cloud AI cheaper.
They can move some inference away from the cloud completely.
As we saw in Will Edge AI Reduce the Need for Giant Data Centers?, a model that fits on a phone, PC, vehicle, or factory device can run locally.
That means some tasks may disappear from the cloud workload.
So model efficiency can pull infrastructure in two directions at once.
Smaller models → less compute per task + more on-device AI
but also:
Cheaper AI → more applications + more frequent AI use
This is why it is dangerous to make a simple forecast such as:
“AI models will become more efficient, therefore data-center demand will fall.”
We need to know what happens to demand too.
The IEA Already Models This Uncertainty
The IEA does not assume there is only one future.
Its 2025 Energy and AI report includes a High Efficiency case in which stronger progress in software, hardware, and infrastructure efficiency reduces global data-center electricity use by more than 15% in 2035 compared with its base case.
That is important.
Efficiency really can reduce the size of the electricity problem.
But even the newer 2026 evidence shows why efficiency cannot be viewed in isolation: use is expanding quickly, and new AI tasks are becoming more computationally intensive.
So the question is not whether efficiency matters.
It clearly does.
The question is whether efficiency improves faster than total AI use expands.
What Should Investors Watch?
If you are watching the AI infrastructure boom, “model size” alone is not enough.
Four numbers become more useful:
- Cost per useful AI task — is inference getting dramatically cheaper?
- Tasks per user — are people asking AI to do more things more often?
- Users and applications — how quickly is AI spreading into new products?
- Compute intensity of new workloads — are agents, reasoning, and video replacing simple text queries?
This changes how we think about companies selling AI infrastructure.
Better chips and smaller models can reduce the hardware needed for one job.
But if lower cost unlocks a much larger market, total demand for accelerators, memory, networking, cooling, and electricity can still grow.
Efficiency is not necessarily the enemy of infrastructure demand.
It can also be one of the things that makes mass adoption possible.
The Number to Remember
If you remember one equation from this article, make it this:
Total Compute = Compute per Task × Number of Tasks
Most headlines focus on the first half.
The future of data centers may depend just as much on the second.
What to Watch Next
There is one more way to rethink the physical cost of data centers.
Instead of asking only how to use less electricity, we can ask what happens to the energy after the computers use it.
Almost all of that electricity eventually becomes heat.
Today, much of that heat is simply rejected into the environment.
But in a cold city with a district-heating network, could that waste heat become useful?
That leads to the next article:
Can Data Center Waste Heat Warm a City?
Key Vocabulary
Small language model (SLM)
A language model designed to use fewer parameters and computing resources than very large frontier models. Smaller models can be useful for specialized, local, or cost-sensitive tasks.
Inference
The process of using a trained AI model to generate an answer, prediction, image, or action.
Quantization
A technique that represents model numbers with lower numerical precision, often reducing memory and computing needs.
Jevons paradox
The possibility that greater efficiency lowers the cost of using a resource enough to increase total consumption.
Rebound effect
The portion of an expected efficiency saving that is offset because lower cost encourages greater use.
Related Articles
- Can We Put AI Data Centers in Space?
- Space Is Cold. So Why Is Cooling a Data Center in Orbit So Hard?
- Will Edge AI Reduce the Need for Giant Data Centers?
- Why Power, Not Chips, May Limit the AI Data Center Boom
- How Much Power Is 100 MW? An AI Data Center Compared with Entire Cities
Sources
- International Energy Agency — Key Questions on Energy and AI: Executive Summary, 2026
- International Energy Agency — Energy and AI: Energy Demand from AI, 2025
- Google Developers Blog — Introducing Gemma 3 270M: The Compact Model for Hyper-Efficient AI, 2025
- Microsoft Azure Blog — One Year of Phi: Small Language Models Making Big Leaps in AI, 2025
- Microsoft Research — Phi-Reasoning: Small and Efficient AI, 2025
- Luccioni, Strubell & Crawford — From Efficiency Gains to Rebound Effects: The Problem of Jevons' Paradox in AI's Environmental Debate, 2025
Published: August 2026 · Sources checked through: August 2026 · Efficiency and energy use vary substantially by model, hardware, prompt length, output length, workload, and deployment architecture.