Smaller AI Models, More AI Use: Why Data Center Demand May Still Grow

Suppose an AI task becomes ten times cheaper to run.

It sounds obvious what should happen next.

We should need fewer GPUs. Less electricity. Less cooling. Maybe fewer giant data centers.

And for that one task, that can be true.

But now change one thing.

Because AI became ten times cheaper, people start using it twenty times more often.

Suddenly, the efficient system can require more total computing than before.

Efficiency tells us how much compute one task needs. Data-center demand also depends on how many tasks we decide to run.

This is the puzzle behind smaller AI models.

Smaller models can reduce cost, power, memory, and cloud dependence. But by making AI useful in more products, for more people, more often, they can also expand total demand.

By the end of this article, you should be able to separate those two forces and judge whether a new efficiency claim is likely to reduce data-center demand—or simply make much more AI affordable.

First, Smaller Does Not Mean Useless

For a while, the AI story sounded simple: bigger model, more computing, better capability.

That still matters at the frontier. The largest and most capable models can require enormous training clusters and expensive inference infrastructure.

But many everyday jobs do not need the largest model available.

A spell checker does not need to solve graduate-level physics.

A factory camera may only need to recognize a few specific faults.

A browser may need a compact model for rewriting, summarizing, or classifying text locally.

This is why small language models, or SLMs, matter. They use fewer parameters and typically need less memory and compute than frontier-scale models.

The important question is not:

Can this small model beat the biggest model at everything?

It is:

Can this smaller model do this particular job well enough?

The Latest Examples Show How Fast the Boundary Is Moving

Google's Gemma 3 270M is one clear example. Google designed the 270-million-parameter model for efficient, task-specific use and reported that an INT4-quantized version used only 0.75% of a Pixel 9 Pro battery during 25 internal test conversations.[2]

That model is not a replacement for every frontier model.

It shows that some useful AI can fit into a very small compute envelope.

Google's newer Gemma 4 12B pushes the other end of the local range. Google says it is small enough to run locally on laptops with 16 GB of VRAM or unified memory, while supporting multimodal inputs.[3]

Microsoft is moving in the same direction. Its current Edge documentation says the prerelease Aion-1.0-Instruct model is significantly smaller, faster, and more efficient than Phi-4-mini and can run even on devices with weaker GPUs or no GPU through CPU inference.[4]

The exact product names will change.

The direction is the important part:

Useful AI is moving into smaller compute budgets.

What Does Better Efficiency Do to a Data Center?

Imagine one AI request needs 10 units of compute.

Engineers improve the model, numerical precision, software, and hardware until the same useful result needs only 1 unit.

If usage stays unchanged:

10× better efficiency → about 90% less compute for that workload

That can mean fewer accelerator-hours, less electricity, less cooling, and lower operating cost.

This part is real.

Efficiency is not an illusion.

The IEA's 2026 update says energy use per individual AI task has been falling at an extraordinary rate. It estimates that software and hardware progress reduced electricity use per task by at least an order of magnitude annually in recent years.[1]

If that were the only thing changing, the electricity story would be simple.

But it is not.

Why Is Total Data-Center Electricity Still Rising?

The same IEA update reports that global data-center electricity use grew about 17% in 2025.

Electricity consumption at AI-focused data centers grew much faster—about 50%.[1]

So two things happened at the same time:

Energy per AI task fell sharply.

Total AI data-center electricity rose sharply.

There is no contradiction once we include demand.

The Four-Variable Equation Behind the Puzzle

Here is the main framework for this article:

Data-center compute ≈ Users × AI tasks per user × Compute per task × Cloud share

Think of each term separately.

  • Users: How many people, businesses, machines, and agents use AI?
  • AI tasks per user: How often is AI called?
  • Compute per task: How expensive is each task?
  • Cloud share: What fraction still runs in a data center rather than on a device or nearby edge system?

Smaller, more efficient models push the third term down.

If they can run locally, they may also push the fourth term down.

But cheaper AI can push the first two terms up.

That is the entire puzzle.

A Simple Example Makes the Rebound Easy to See

Suppose a company runs 1 million AI tasks per day.

Each task needs 10 compute units.

Total compute is:

1 million × 10 = 10 million compute units

Now efficiency improves tenfold.

Each task needs only 1 unit.

If usage stays the same:

1 million × 1 = 1 million compute units

Demand falls by 90%.

Now imagine cheaper AI makes new features economical and daily usage rises to 15 million tasks.

15 million × 1 = 15 million compute units

The compute needed for one task fell by 90%.

Total compute still rose by 50% compared with the original system.

These are invented numbers for teaching, not a forecast.

They show why an efficiency headline cannot tell us total infrastructure demand by itself.

Is This the Jevons Paradox?

Sometimes people call this the Jevons paradox.

The idea comes from economics: if efficiency makes a resource cheaper to use, demand can rise enough that total resource consumption increases rather than falls.

For AI, researchers have warned that the same rebound logic may matter. A 2025 FAccT paper by Alexandra Sasha Luccioni, Emma Strubell, and Kate Crawford argues that efficiency gains alone do not guarantee lower environmental impact because lower cost can stimulate more use.[7]

But we should use the term carefully.

Jevons paradox is not a law saying that every efficiency improvement must increase total consumption.

Sometimes efficiency reduces total demand.

Sometimes additional usage offsets only part of the saving. That is usually called a rebound effect.

And sometimes usage grows so strongly that total demand rises above the original level.

So the useful question is not:

Is Jevons paradox happening?

It is:

How large is the usage rebound compared with the efficiency gain?

AI efficiency rebound loop showing lower cost per task leading to more applications, users, and AI calls

Figure 1. Efficiency lowers the cost of one task. Lower cost can expand the number of tasks. Total compute depends on both.

Why AI May Be Especially Sensitive to Rebound

AI has a special feature: new demand can appear very quickly.

A cheaper model does not only make an existing chatbot cheaper.

It can put AI into every email, document, camera, browser tab, industrial machine, vehicle, customer-support flow, and software agent.

And one visible user action may hide many model calls.

An agent asked to “plan a trip” might search, compare, summarize, check prices, revise its plan, call tools, and verify the result.

The user sees one request.

The infrastructure may see dozens or hundreds of inference steps.

The IEA highlights this change directly. It says reasoning, video generation, and agentic tasks can use hundreds or thousands of times more electricity per query than simple text generation.[1]

This is why task intensity matters as much as model size.

Small Models Can Pull Data-Center Demand in Two Directions

There is another force from the previous article in this series.

In Will Edge AI Reduce the Need for Giant Data Centers?, we saw that a model small enough to run on a phone, PC, vehicle, or factory system can remove some inference calls from the cloud.

So smaller models can reduce data-center demand in two ways:

  • less compute is needed for one task, and
  • some tasks can leave the cloud entirely.

But the same lower cost can also increase demand:

  • more users adopt AI,
  • each user calls AI more often,
  • developers place AI inside more products, and
  • agents create longer chains of model calls.

This is why “small models mean fewer data centers” is too simple.

Efficiency Still Matters—A Lot

It would be equally wrong to conclude that efficiency is useless because of rebound.

The IEA's 2025 Energy and AI scenarios show why.

In its High Efficiency case, stronger improvements in software, hardware, and data-center infrastructure reduce global data-center electricity demand by more than 15% in 2035 compared with the base case.[6]

So efficiency can materially reduce infrastructure pressure.

The uncertainty comes from the other side of the equation: how fast AI use expands and how compute-intensive the new workloads become.

The better question is therefore:

Is efficiency improving faster than useful AI demand is expanding?

What Does “Smaller” Actually Mean?

Model parameter count is useful, but it is not the whole story.

Engineers can reduce inference cost in several ways:

  • Smaller models: fewer parameters to store and process.
  • Quantization: represent numbers with lower precision so models use less memory and computation.
  • Distillation: train a smaller model to reproduce useful behavior from a larger teacher model.
  • Specialization: use a model designed for one narrow task instead of a general model for everything.
  • Cascades and routing: let a small model handle easy requests and send only difficult ones to a larger model.
  • Faster inference software and hardware: improve how efficiently the same model runs.

Google Research has explored this routing idea directly. Its speculative-cascade work combines smaller and larger models so inexpensive requests can be handled efficiently while harder cases defer to a more capable model.[5]

This gives us another useful rule:

Do not use the biggest model by default. Use enough model for the job.

Six Numbers Matter More Than Model Size Alone

When you see a headline saying a new small model will cut AI infrastructure demand, check these six numbers.

  1. Compute per useful task: How much cheaper is the same useful result?
  2. Quality at the target task: Does the smaller model actually do the job well enough?
  3. Tasks per user: Does lower cost make people use AI much more often?
  4. Number of users and applications: Is AI spreading into new products and workflows?
  5. Cloud share: Does the smaller model run locally, or is it still served from a data center?
  6. Task intensity: Are new workloads simple text, or are they reasoning, agents, video, and long-context work?

Those six numbers tell us much more than parameter count alone.

So Will Smaller AI Models Shrink the Data-Center Boom?

Now we can return to the opening puzzle.

Yes, smaller and more efficient models can reduce compute, electricity, and cooling for a fixed amount of work.

They can also move some inference onto devices and away from cloud infrastructure.

But they can make AI cheap enough to appear in many more places.

And the new AI workloads that remain in the cloud may become much more demanding.

So the direction of total data-center demand depends on all four terms:

Users × AI tasks per user × Compute per task × Cloud share

Smaller models push some terms down.

Cheaper and more useful AI can push others up.

That is why efficiency can slow the data-center boom without necessarily ending it.

What to Watch Next

If you want to know whether efficiency is actually reducing infrastructure pressure, watch these signals:

  • energy or compute per useful AI task,
  • inference cost per task,
  • AI tasks per user,
  • the number of AI-enabled products and agents,
  • the share of inference moving to devices,
  • the growth of reasoning, video, agentic, and long-context workloads.

The next article asks what happens to the energy after all that computation is finished.

Can Data Center Waste Heat Warm a City?

Key Terms

small language model (SLM)
A language model designed to use fewer parameters and computing resources than very large frontier models. The useful comparison is whether it can perform the target job at acceptable quality.

inference
Using a trained model to generate an answer, prediction, image, action, or other output.

quantization
Representing model values with lower numerical precision to reduce memory use and computation.

distillation
Training a smaller model to reproduce useful behavior learned from a larger model.

Jevons paradox
A situation in which efficiency makes a resource cheaper to use and total consumption rises enough to offset—or exceed—the expected savings.

rebound effect
The part of an expected efficiency saving that is lost because lower cost encourages more use.

task intensity
How much computation one AI task requires. A simple text rewrite and a multi-step agent workflow can have very different intensity.

Related Articles

Sources

  1. International Energy Agency — Key Questions on Energy and AI: Executive Summary, 2026.
  2. Google Developers Blog — Introducing Gemma 3 270M.
  3. Google Developers Blog — Gemma 4 12B: The Developer Guide, June 3, 2026.
  4. Microsoft Learn — Phi Silica / transition to Aion Instruct, current October 2026 documentation.
  5. Google Research — Speculative cascades: A hybrid approach for smarter, faster LLM inference.
  6. International Energy Agency — Energy and AI: Energy Demand from AI, 2025.
  7. Luccioni, Strubell & Crawford — From Efficiency Gains to Rebound Effects: The Problem of Jevons' Paradox in AI's Polarized Environmental Debate, FAccT 2025.

Updated: October 2, 2026 · Sources checked through: October 2, 2026 · The equations and numerical examples are teaching tools, not forecasts of future AI usage or data-center demand.