AI/ML · 10 min read · Sep 17, 2026

AI Is Getting More Efficient. So Why Does It Need More Computing Power?

AI models are becoming more efficient, but that is not reducing the industry's appetite for computing power. As inference takes a larger share of AI workloads, data centers, accelerators, electricity and cooling infrastructure are becoming central to the next phase of AI.

AI Is Getting More Efficient. So Why Does It Need More Computing Power?

One of the assumptions surrounding the next phase of artificial intelligence is that better models and more efficient chips will eventually reduce the amount of computing power AI needs.

There is good reason to expect efficiency gains. AI hardware is becoming faster. Software is getting better at squeezing more work out of each chip. Smaller models can perform tasks that once required much larger systems.

Yet the overall demand for computing power is moving in the opposite direction.

The reason is simple: AI is being used more often, for more complicated tasks, and for longer periods of time.

Deloitte's 2026 Technology, Media & Telecommunications Predictions puts this trend at the center of its second prediction. The firm expects AI inference—the process of running an already-trained model to produce an answer, prediction or generated output, to account for about two-thirds of all AI computing in 2026. Deloitte also expects most inference to remain in data centers and enterprise servers rather than shifting predominantly to inexpensive, low-power edge devices.

That has major consequences for Nvidia and its competitors, cloud providers, data-center operators and the electricity infrastructure supporting them.

Training gets the attention. Inference may become the bigger workload

The public conversation around AI computing has often focused on training.

Training is the process in which a model learns patterns from enormous amounts of data. It can require thousands of accelerators operating simultaneously for extended periods.

Inference is different.

Every time someone asks an AI chatbot a question, generates an image, uses an AI coding assistant or interacts with an AI agent, a trained model has to perform inference.

One user making one request may require relatively little computing power.

Billions of users making requests continuously is another matter.

And AI applications are becoming more computationally demanding.

Reasoning models can spend additional compute working through a problem before producing an answer. AI agents can perform multiple model calls while completing a task. Longer context windows mean more information can need to be processed. Generative video can require substantially more computation than a short text response.

The industry's transition from training a relatively small number of frontier models to serving those models to millions or billions of users therefore changes the economics of AI infrastructure.

Deloitte expects this transition to push inference toward roughly two-thirds of total AI compute in 2026.

Better chips are not cancelling out rising demand

There is an important distinction between efficiency per AI task and total AI computing demand.

A newer accelerator might produce substantially more AI output for every watt of electricity than an older generation.

That is good news.

But if the number of AI tasks grows faster than efficiency improves, total energy consumption can still rise.

The International Energy Agency's latest analysis illustrates the problem.

The IEA says electricity consumption from data centers increased by 17% in 2025, while electricity use by AI-focused data centers grew even faster. At the same time, power consumption per AI task is declining rapidly because of improvements in hardware and software. Yet the IEA still expects global data-center electricity consumption to roughly double from 485 terawatt-hours in 2025 to around 950 TWh in 2030. Electricity consumption from AI-focused data centers is projected to triple during that period.

In other words, efficiency is improving.

Usage is growing faster.

This is sometimes described as a rebound effect: making something cheaper or more efficient can encourage people to use more of it.

AI appears to be experiencing a version of that dynamic.

Nvidia is designing around the inference economy

Nvidia remains the dominant supplier of the accelerators used throughout much of the AI infrastructure stack, but the company's pitch is increasingly shifting toward the economics of inference.

That means operators care about more than how powerful a chip is.

They want to know how many useful tokens it can generate per watt, how quickly it can respond, how many users it can serve and how much each request costs.

Nvidia's Blackwell platform is built around that calculation.

Nvidia says its GB300 NVL72 system can deliver up to 50 times the throughput per megawatt and up to 35 times lower cost per token than Hopper for certain low-latency agentic workloads, based on SemiAnalysis InferenceX benchmarks. Those are Nvidia-reported comparisons, rather than independent measurements applicable to every workload.

The significance is less about any single benchmark than about what the benchmark reveals about the industry.

AI infrastructure is increasingly being measured in tokens per dollar and tokens per watt.

That is a different conversation from the early AI boom, when raw accelerator performance and model-training capacity dominated attention.

AMD and other chipmakers want part of the inference market

Nvidia is not the only company designing hardware around this opportunity.

AMD has been expanding its Instinct accelerator portfolio and competing directly for large AI deployments.

In September 2026, AMD submitted results across six model families in the MLPerf Inference 6.1 benchmark, using its MI355X and MI350X accelerators as well as the MI350P PCIe card. The workloads included language, reasoning, text-to-video and recommendation inference.

AMD has also been working on lower-precision computing and distributed inference approaches designed to improve performance and cost efficiency.

Google is another important competitor, using its internally developed Tensor Processing Units in its own AI infrastructure and cloud services.

Google reported in April that its seventh-generation Ironwood TPU achieved an approximately 3.7-fold improvement in compute carbon intensity compared with its previous performance-optimized TPU generation.

The competitive landscape therefore extends beyond a simple Nvidia-versus-AMD chip race.

Cloud providers are developing their own silicon. Startups are building specialized inference processors. Hyperscalers are designing complete systems around their preferred accelerators.

The common objective is to produce more AI output from every unit of computing infrastructure.

Data centers are becoming an energy infrastructure problem

The computing challenge eventually becomes an electricity challenge.

A modern AI data center does not only require processors.

It requires networking, memory, storage, cooling, backup power and the electrical infrastructure capable of keeping all of those systems operating continuously.

AI accelerators are particularly power-dense, which creates another problem: heat.

As computing density rises, cooling systems have to remove more heat from the same physical space. That can make the availability of electricity only one part of the constraint.

The grid connection, transformers, substations, cooling systems and physical location of the facility can all become limiting factors.

The IEA estimates that data-center electricity consumption could reach roughly 950 TWh globally by 2030, equivalent to around 3% of global electricity demand. It also warns that bottlenecks involving grid connections, transformers, gas turbines, advanced chips and other infrastructure are already affecting the pace at which new capacity can come online.

Deloitte's own analysis has similarly highlighted the growing connection between AI infrastructure and electricity infrastructure.

That makes AI expansion partly an energy-planning problem.

The edge is still important, but it may not absorb most AI inference

There is another part of Deloitte's prediction that deserves attention.

The industry has spent years discussing edge AI: running models directly on phones, PCs, vehicles, cameras and other devices instead of sending every request to a distant data center.

There are obvious advantages.

Local inference can reduce latency. It can improve privacy. It can reduce the amount of information that needs to travel across networks.

And AI-capable consumer devices are becoming more common.

But Deloitte does not expect that to displace data-center inference at scale in 2026.

Instead, the likely architecture is more complicated.

Some AI workloads will run locally.

Others will run in enterprise infrastructure.

The most computationally intensive workloads will continue to depend on large centralized systems.

A smartphone might handle speech recognition or a small generative model locally while sending a difficult reasoning task to a cloud model.

The future is therefore likely to be hybrid rather than entirely centralized or entirely decentralized.

AI infrastructure spending is becoming enormous

The scale of investment required to support this transition is difficult to ignore.

Deloitte expects global capital expenditure related to AI data centers to reach roughly $400 billion to $450 billion in 2026.

The IEA provides another indication of the scale. Capital expenditure by five major technology companies exceeded $400 billion in 2025 and is expected to rise by another 75% in 2026, according to its April 2026 analysis.

That spending is not limited to servers.

It covers data-center construction, networking, power infrastructure, cooling and other equipment needed to operate increasingly dense AI clusters.

It also creates a financial question.

AI companies need enough demand to justify this infrastructure.

If utilization remains low, expensive accelerators can sit idle and push up the cost of each generated token. If demand grows rapidly, infrastructure operators have an incentive to keep building.

That makes utilization one of the most important variables in the AI infrastructure economy.

What this means for cloud providers and AI companies

The economics of AI services will increasingly depend on infrastructure efficiency.

An AI company may have an excellent model, but if every answer is expensive to generate, scaling the product can become difficult.

This is why model compression, quantization, caching, efficient inference software and specialized accelerators are becoming strategic technologies.

The objective is not simply to make AI smarter.

It is to make AI cheap enough and fast enough to use continuously.

That distinction could influence which AI products become commercially sustainable.

It also helps explain why cloud providers are investing so aggressively in custom silicon and specialized infrastructure.

The winning architecture for an AI workload may not always be the chip with the highest theoretical performance.

It may be the system that produces the required result at the lowest cost within the available power and cooling envelope.

The energy question is becoming impossible to separate from the AI question

For years, discussions about AI progress focused primarily on algorithms and model capabilities.

That is changing.

The next constraints may be physical.

Can enough chips be manufactured?

Can data centers secure electricity?

Can utilities build transmission and generation capacity quickly enough?

Can operators cool increasingly dense systems?

Can companies earn enough revenue from AI services to justify the infrastructure spending?

And can the industry continue increasing AI usage while reducing the environmental impact of each individual task?

There are reasons for optimism on efficiency.

The IEA says energy efficiency per AI task is improving rapidly. Google reports significant improvements in the carbon efficiency of its latest TPU generation. Nvidia and AMD are competing to increase performance per watt and lower inference costs.

But those improvements do not automatically mean total consumption will fall.

If AI becomes cheaper, faster and more useful, people are likely to use more of it.

What to watch next

Deloitte's prediction is ultimately less about AI chips becoming bigger and more about AI becoming infrastructure.

When inference accounts for a growing share of computing, every interaction with an AI system has a physical cost somewhere: processors, memory, networking, electricity and cooling.

The industry is working aggressively to reduce that cost.

But efficiency alone may not be enough to reverse the growth in demand.

The next phase of AI will therefore be shaped not just by who builds the smartest model, but by who can operate that intelligence at scale.

That puts Nvidia, AMD, Google and other accelerator developers in a race over performance per watt. It puts cloud providers in a race over infrastructure economics. And it puts utilities and governments in the increasingly important position of determining whether enough power can reach the machines that run the AI economy.

The central question for 2026 is no longer whether AI will need more computing.

It is how much useful intelligence the industry can produce from every watt, chip and data-center rack it builds.