You buy a new laptop.
The spec sheet says it has a CPU, a GPU—and now an NPU.
Then you open Task Manager, run an AI app, and the NPU seems to do almost nothing.
So what is it actually for?
An NPU is not a replacement for the CPU or GPU. It is a specialized engine for running supported AI inference efficiently—especially when power, heat and continuous local operation matter.
First, CPU vs. GPU vs. NPU in Plain English
Imagine a computer has three kinds of workers.
CPU: the flexible generalist. It handles operating-system logic, application control, branching, small tasks and almost anything that needs versatility.
GPU: the large parallel team. It is very good when thousands of similar calculations can run at once, which is why GPUs became central to graphics, AI training and high-throughput generative AI.
NPU: the efficient specialist. It is built around the repetitive tensor and neural-network operations used by AI inference, with a strong focus on doing that work with less power.
A modern device can contain all three because the best processor depends on the task.
Quick Answer
NPU stands for Neural Processing Unit.
It is specialized hardware designed to accelerate neural-network inference.
Inference means using a trained model to produce an output: recognize a face, transcribe speech, summarize text, enhance a camera image, classify an object or generate tokens.
Microsoft's current Windows ML documentation describes the three roles this way: NPU for battery-efficient sustained on-device inference, GPU for high-throughput image, video and generative AI, and CPU as a universal fallback.[1]
Original Asset 1: Right Workload → Right Processor
general logic / control
→ CPU
large parallel graphics / high-throughput GenAI
→ GPU
supported sustained inference under a tight power budget
→ NPU
These are not hard borders.
The same model may run on more than one kind of processor.
Why Build an NPU If the GPU Can Already Run AI?
Because maximum speed is not always the goal.
Consider an AI feature that listens for speech, cleans up a video call, watches a camera feed or runs a small local assistant throughout the day.
If a powerful GPU wakes up for every small inference task, the device may use more power and produce more heat than necessary.
An NPU can be more useful when the task is:
- repetitive,
- supported by the NPU runtime,
- small or medium enough for the device,
- latency-sensitive,
- and expected to run for long periods.
Why Your NPU Can Sit at 0%
This is one of the most common points of confusion.
Owning an NPU does not mean every AI application automatically uses it.
Microsoft explicitly says an NPU is a hardware resource that software must be programmed to use.[3]
The application needs a compatible runtime and a hardware-specific execution path.
In Windows ML, these paths are provided through execution providers. An execution provider connects the model runtime to a specific CPU, GPU or NPU backend and handles hardware-specific optimization.[2]
If the app does not support the NPU—or if the model contains operations the NPU backend cannot run—the workload may use the GPU or CPU instead.
An NPU can be physically present but practically invisible until the software stack knows how to use it.
Original Asset 2: The NPU Usefulness Gate
Before expecting an NPU to help, ask:
- Is this inference? NPUs are usually optimized more for inference than heavy model training.
- Does the runtime support the model?
- Does the NPU backend support the model's operators?
- Does the model fit the available memory and device limits?
- Is the task repeated or sustained long enough for efficiency to matter?
- Does the application actually select the NPU path?
- Are battery life, heat, privacy or offline use important?
- Would the GPU or cloud still be better for this job?
If several answers are no, the NPU may contribute little even if the spec sheet advertises a large TOPS number.
What Is TOPS?
TOPS means trillions of operations per second.
It is a peak compute-rate metric often used for AI accelerators.
Microsoft currently uses an NPU capable of 40+ TOPS as part of the Copilot+ PC hardware definition.[3]
That number is useful for defining a platform class.
It is not the same thing as saying:
“This laptop will generate twice as many tokens per second as a laptop with half the TOPS.”
Original Asset 3: TOPS Reality Check
peak TOPS
≠
real application performance
Real AI performance also depends on:
- numerical precision,
- model architecture,
- which operators are supported,
- memory bandwidth,
- model size,
- runtime and compiler quality,
- data movement,
- batch size,
- and thermal / power limits.
AMD explicitly notes that its TOPS figures describe a maximum under optimal conditions and that actual results can vary with system configuration, AI model and software version.[6]
This is why TOPS should be treated as one specification, not a complete benchmark.
Can an NPU Run a Local LLM?
Yes, some NPUs can run language models locally.
But “can run” and “is the best processor for the entire model” are different questions.
Large language-model inference has different phases.
Prefill processes the input prompt and creates internal state.
Decode generates output tokens one by one.
Those phases can stress hardware differently.
A Useful 2026 Example: One LLM, Two Processors
AMD's current Ryzen AI software provides a particularly useful example.
AMD describes an on-device LLM path where the NPU handles the compute-intensive prefill phase while the integrated GPU handles the memory-bound decode phase.[5]
That is important because it breaks the simple “NPU versus GPU” framing.
The better architecture can be:
prompt processing
→ NPU
token-by-token generation
→ integrated GPU
One AI model can use more than one processor because different parts of the workload have different bottlenecks.
Original Asset 4: Hardware Routing Map
A practical device might route work like this:
| Workload | Likely good fit | Why |
|---|---|---|
| App logic / fallback | CPU | Flexible, universal |
| Large image / video / GenAI workload | GPU | High parallel throughput |
| Background vision / audio / supported local inference | NPU | Power-efficient sustained execution |
| Complex local LLM | NPU + GPU + CPU | Different phases/operators may fit different engines |
| Very large frontier model | Cloud / data center | Local memory and compute may be insufficient |
The Operating System Is Becoming an AI Traffic Controller
The long-term direction is not for users to manually choose a processor every time.
Windows ML exposes hardware-specific execution providers for Intel, AMD, Qualcomm and NVIDIA devices and gives applications a common inference framework across CPU, GPU and NPU.[1][2]
Apple follows a similar system idea. Core ML can use the CPU, GPU and Neural Engine while optimizing for device performance and power consumption.[8]
So the broader trend is:
AI hardware is becoming heterogeneous, and software increasingly decides where each part of the work should run.
Why NPU Performance Is Often a Memory Problem Too
An NPU can perform huge numbers of arithmetic operations.
But the model's weights and intermediate data still have to reach the compute units.
If data movement cannot keep up, more theoretical math capacity does not help.
This is especially important for local generative AI, where model size and memory bandwidth can strongly affect responsiveness.
That is another reason two NPUs with similar TOPS can deliver different real experiences.
What Kind of Tasks Fit an NPU Well?
Good candidates often include:
- speech recognition,
- noise suppression,
- camera enhancement,
- image classification,
- OCR and document analysis,
- small local language models,
- embedding generation,
- always-on sensor or vision inference,
- and other supported background AI features.
The important word is supported.
A theoretically suitable workload still needs a working runtime and model implementation.
What About Phones?
The same idea exists outside PCs, even when vendors use different names.
Qualcomm calls its specialized engine the Hexagon NPU. In September 2026, Qualcomm described a new NPU architecture aimed at increasingly complex on-device agentic AI, with a transformer-focused accelerator, larger shared memory and broader precision support.[7]
Apple uses the name Neural Engine rather than NPU.
These names differ, but the architectural idea is similar: add specialized hardware for efficient local machine-learning inference.
NPU Does Not Mean “Everything Stays Local”
On-device AI has important advantages:
- lower network latency,
- offline operation,
- less data leaving the device,
- lower cloud usage for some tasks,
- and potentially lower power for sustained supported inference.
But local hardware has limits.
A device cannot always hold or run the largest model efficiently.
Cloud systems can offer larger models, more memory, faster server accelerators and continuously updated capabilities.
So many products will use both.
Original Asset 5: The Local-vs-Cloud Cost Path
Local NPU path:
device purchase
→ local inference
→ no per-request cloud accelerator charge
→ lower network dependence
→ privacy / offline benefits
→ limited by local memory, model and software
Cloud path:
request leaves device
→ network
→ data-center accelerator
→ larger / more capable model
→ server compute cost + network dependence
The best architecture may route easy, frequent and private work locally while sending harder tasks to the cloud.
Why This Matters for Robots and Physical AI
A robot, drone or vehicle may need to respond before a cloud round trip is practical.
Local AI hardware can help with perception, audio, sensor interpretation or smaller policy/inference workloads.
But the same routing principle applies: real-time control may stay on CPU or microcontroller, high-throughput perception may use GPU, and supported efficient inference may use NPU.
Physical AI therefore strengthens the case for heterogeneous compute rather than one universal AI processor.
Original Asset 6: The NPU Buying Checklist
If you are buying a laptop, do not compare only TOPS.
Ask:
- Which AI applications do I actually use?
- Do those apps support this NPU?
- How strong is the GPU?
- How much RAM does the system have?
- What is the memory bandwidth?
- Do I care about battery life and fan noise?
- Do I want local/offline AI?
- Will I run large generative models, where GPU and memory may matter more?
For many buyers, software support and memory can matter more than a small difference in advertised NPU TOPS.
What to Watch Next
- Software support. Do more mainstream apps use NPUs automatically?
- Runtime orchestration. Does CPU/GPU/NPU routing become invisible to users?
- Local agents. Do always-on and agentic workloads make sustained NPU efficiency more valuable?
- Memory. Do faster unified-memory systems unlock better local GenAI?
- Hybrid inference. Does splitting prefill, decode and other operators across processors become standard?
- Benchmarks. Do buyers move from TOPS toward model-level latency, tokens/s and performance-per-watt?
- Cloud balance. Which tasks remain local and which continue to use frontier cloud models?
The Simple Idea to Remember
An NPU is not “the new GPU.”
It is another specialized compute engine.
The useful question is not “CPU, GPU or NPU?” It is “Which processor should run this part of the AI workload at the right speed, power and cost?”
Key Vocabulary
NPU
Neural Processing Unit: specialized hardware for efficient neural-network inference.
inference
Using a trained model to produce an output.
TOPS
Trillions of operations per second, a peak AI-compute metric.
execution provider
A software component that lets an AI runtime use a particular CPU, GPU or NPU backend.
prefill
The phase in an LLM that processes the input prompt before generating output tokens.
decode
The token-by-token generation phase of an LLM.
quantization
Representing model values with lower numerical precision to reduce memory and compute requirements.
Neural Engine
Apple's name for its specialized on-device neural-network accelerator.
Read the AI Hardware Full Stack Series
- Why AI Chips Need Advanced Packaging, Not Just Smaller Transistors
- What Is a Chiplet? Why AI Chips Are Splitting Into Specialized Dies
- Why Silicon Photonics Could Become the Next AI Data Center Bottleneck
- What Is an NPU? Why AI Does Not Have to Run Only on GPUs
- The Hidden Chips Behind AI Power: Why Power Semiconductors Matter
Related Articles
- Training vs. Inference: Why AI Needs Different Infrastructure for Each
- GPU vs. HBM: Why AI Needs Both Compute and Memory
- What Is Physical AI?
Sources
- Microsoft — Windows ML overview, checked October 4, 2026.
- Microsoft — Accelerate AI models with Windows ML, checked October 4, 2026.
- Microsoft — Copilot+ PC developer guide / NPU access, checked October 4, 2026.
- Microsoft — Choose your Windows AI solution, checked October 4, 2026.
- AMD — Hybrid NPU/GPU on-device LLM inference, May 21, 2026.
- AMD — Ryzen AI Max PRO 400 Series and TOPS caveat, May 2026.
- Qualcomm — Hexagon NPU architecture for agentic AI, September 10, 2026.
- Apple — Core ML, checked October 4, 2026.
- Apple — Generative models / Core AI, checked October 4, 2026.
- Intel — NPU exploration and explicit workload offload, April 28, 2026.
Sources checked through October 4, 2026. NPU is not a universal architecture or naming standard. TOPS is a peak compute metric, not a complete predictor of application performance. Vendor performance claims are attributed to their publishers.