Our Review Methodology: Every tech review on AIGadgeTech synthesizes analysis from leading publications including TechRadar, CNET, TechCrunch, The Verge, and Wired. We cross-reference professional testing data with real user experiences from Reddit, X (Twitter), and tech forums to give you the complete picture—not just manufacturer claims. Our comparison approach identifies consensus across sources and highlights where reviews conflict, so you can make informed decisions based on comprehensive data.

What Is an AI NPU in Phones? Plain-English Guide (2026)

What Is an AI NPU in Phones? Plain-English Guide (2026)

Illustration for what is an ai npu in phones

An AI NPU (neural processing unit) is a dedicated block of silicon inside your phone's main chip that runs artificial-intelligence calculations quickly while using very little battery. Instead of forcing the main processor (CPU) or graphics chip (GPU) to grind through AI math, the NPU handles it at a fraction of the power cost. It's what lets a modern phone blur your video-call background, transcribe speech, or run a language model offline without getting hot. Apple calls its version the Neural Engine, Qualcomm calls its Hexagon, Google's Tensor chip carries a TPU, and MediaTek uses the name APU — different labels, same job.

Two years ago, roughly half of flagship phones shipped with a dedicated NPU. In 2026, it's over 80%, and top flagship NPUs now reach around 75 TOPS — enough to run a 7-billion-parameter language model entirely on the device, according to Fastio's 2026 AI-phone roundup. That shift is why "NPU" suddenly appears in every phone launch.

Illustration of a smartphone processor highlighting the NPU block alongside the CPU and GPU

About the author

Chester Takau is an independent tech reviewer who synthesises professional testing data with real user experience to cut through marketing claims.

What an NPU actually does that the CPU and GPU can't

Almost every phone AI task — recognizing a face, transcribing your voice, stacking a night-mode photo, generating a text reply — boils down to the same thing: enormous numbers of small multiply-and-add calculations run at low precision. Your phone already had two ways to do that math, and both have problems.

The CPU is a handful of powerful cores built for sequential logic. It makes decisions fast, but asking it to run millions of parallel multiplications is like asking one accountant to do the work of ten thousand calculators. The GPU has those ten thousand calculators — thousands of small parallel cores — but it draws serious power and builds up heat when it runs AI math for more than a few seconds.

The NPU is a third path: silicon shaped specifically for matrix math at reduced precision (INT8, increasingly INT4), placed close to memory so data barely has to travel. Testing cited by Samsung Semiconductor and independent reviewers puts the same AI workload at roughly 5–10W on an NPU versus 30–40W on a GPU. That gap is the whole point.

Chip Built for AI weakness Typical phone jobs
CPU General logic, one task at a time done very fast Slow and power-hungry at massive parallel math Apps, the OS, web browsing
GPU Parallel graphics rendering High power draw and heat under sustained AI load Games, video, some photo effects
NPU Low-precision matrix math for AI models Only runs supported model types; needs software written for it Live translation, transcription, scene detection, on-device language models

This efficiency gap answers a question readers ask a lot: yes, the NPU is a big reason your battery survives camera AI features. Night-mode stacking, subject detection, and real-time HDR all run constantly while you shoot. On a GPU, that sustained load would drain the battery and heat the phone; on an NPU, it ticks along in the background.

TechQuickie's explainer walks through how Apple, Google, and Qualcomm all ended up putting this silicon in your pocket:

Neural Engine vs Hexagon vs Tensor TPU vs APU: the name decoder

Half the confusion around phone NPUs is branding. Every chipmaker builds the same kind of block and gives it a different name:

Vendor NPU brand Where you'll find it 2026 notes
Apple Neural Engine A-series chips in iPhones Decides which Apple Intelligence tier your iPhone gets
Qualcomm Hexagon NPU Snapdragon 8-series (Galaxy S and most Android flagships) Snapdragon 8 Elite Gen 5 claims a 37% faster NPU that can run 56+ AI models in under 5ms (Gizmochina)
Google TPU (Tensor Processing Unit) Tensor chips in Pixel phones Runs Gemini Nano on-device
MediaTek APU (NPU 990) Dimensity chips in OPPO and vivo flagships Dimensity 9500 is the industry's first compute-in-memory NPU, claiming 100% faster 3B-parameter LLM output; OPPO's Find X9 shipped it first (MediaTek)

Samsung sits in both camps: Galaxy flagships in most regions use Snapdragon's Hexagon, while Exynos models carry Samsung's own NPU. If a spec sheet mentions any of these names — or a TOPS figure — the phone has dedicated AI silicon.

Does your phone actually have an NPU, or is that just marketing?

If you bought a flagship from 2023 onward, the answer is almost certainly yes. To check, find your chip name (Settings → About Phone, then search the model), and look at its spec sheet for "Neural Engine," "Hexagon," "APU," "TPU," or a TOPS rating.

Budget and mid-range phones are murkier. Many include a modest NPU; some have none worth mentioning. When a cheap phone advertises itself as an "AI phone," that usually means the AI happens in the cloud — your data goes to a server, and the phone's own silicon does little of the work. Owning an NPU also says nothing about whether your apps actually use it, which is a separate problem we'll get to.

Do TOPS numbers actually mean anything?

TOPS stands for trillions of operations per second, and it's the number vendors slap on every NPU launch. Treat it with suspicion. The Register's coverage of the NPU debate makes the case well: TOPS is a theoretical peak measured under lab conditions, vendors inflate it by quoting lower-precision math (an INT4 figure looks twice as big as INT8), and it says little about how fast a real feature feels. One chip-industry CEO, Quadric's, went on record calling NPU marketing a "scam" — a rare insider saying the quiet part aloud.

The fair reading: TOPS works like a horsepower figure. It tells you the rough class of the chip, not the experience. Memory bandwidth and — more than anything — whether software is written to use the NPU matter just as much. The same TOPS one-upmanship is now playing out in laptops; we untangle that version in our guide to AI laptop features, from NPUs and TOPS to Copilot+.

Why your older flagship got locked out of the new AI features

This is where the NPU stopped being a spec bullet and started being a gatekeeper. In June 2026, MacRumors reported that Apple's most powerful on-device AI model requires an iPhone 17 Pro or iPhone Air — locking out even some current-generation iPhones, while baseline Apple Intelligence still demands at least an iPhone 15 Pro. Google has done the same thing: its Gemini Intelligence tier requires a 2026 flagship chip, 12GB+ of RAM, and Gemini Nano v3, explicitly excluding Nano v2 devices like the Galaxy S25 and Pixel 9 series (Aprenderhub). Google's naming moves fast — it has already formally unveiled Gemini Nano 4 for rollout through 2026.

What this guide adds: most coverage reports these exclusions as isolated news. The table below maps them side by side, so you can see exactly what each NPU tier costs you in features.

Phone / tier NPU status What you get What you're locked out of
iPhone 17 Pro / iPhone Air Latest Neural Engine + memory bar Apple's most powerful on-device AI model
iPhone 15 Pro through iPhone 16 Older Neural Engine Baseline Apple Intelligence The top-tier on-device model
2026 Android flagships (Snapdragon 8 Elite Gen 5, Dimensity 9500) Gemini Nano v3-capable Gemini Intelligence tier
Galaxy S25, Pixel 9 series Gemini Nano v2 only Earlier on-device Gemini features Gemini Intelligence tier

Critics read this as a planned-obsolescence lever: a chip block you can't upgrade now decides which software features you're allowed to run. Whether the limits are genuine engineering or upgrade-cycle pressure, the practical effect is the same — NPU generation has become a buying decision, not a footnote.

Is Gemini Nano actually running on your phone, or in the cloud?

Gemini Nano is built to run locally on the NPU — no server round trip for the tasks it supports, which is why it works offline and why Google can call it private. But "Gemini" on a phone is a hybrid: heavier requests get handed to cloud models, and the phone decides which path to take. The honest answer is that on-device handles the light, fast, private jobs, while the cloud still does the heavy lifting.

The same on-device logic is spreading across other gadgets. It's the reason AI security cameras can tell a person from a passing car without uploading your footage. Smart speakers sit at the opposite extreme — most still send everything past the wake word to the cloud, which is worth knowing if you're comparing the current field in our best AI smart speakers ranking.

Will third-party apps ever use the NPU?

Mostly they don't yet, and the industry has itself to blame. Android Authority's critique lays it out: eight years in, Qualcomm, Apple, MediaTek, and Samsung each built proprietary, incompatible NPU software stacks with no CUDA-equivalent, so an app developer would have to write the same feature four different ways. Google's LiteRT (its 2024 rebrand of TensorFlow Lite) is the first serious attempt at a unifying runtime. On iPhones, apps that use Apple's Core ML framework get routed to the Neural Engine automatically, which is why iOS apps quietly use the NPU more often than Android apps do. For now, the features that lean hardest on your NPU are still the phone maker's own.

Should you buy a phone for its NPU?

If you use camera AI, live translation, transcription, or assistant features daily, NPU quality matters and the efficiency gains are real. If you don't, a bigger TOPS number shouldn't drive your purchase — but the gating trend changes the math. Features arriving through 2026 and 2027 increasingly require current-generation NPUs, so a stronger one buys you feature longevity even if you ignore AI today.

The industry is betting in that direction. Qualcomm is reportedly exploring a standalone NPU paired with custom 3D DRAM (~40 TOPS) for devices landing late 2026 or early 2027 (Intelligent Living), and OpenAI is reportedly fast-tracking its own phone with a dual-NPU architecture for production as early as the first half of 2027 (9to5Mac). NPUs are moving from "extra block on the chip" toward the center of phone design.

Frequently asked questions

Is the NPU the reason my battery lasts longer when I use camera AI features?

Yes, largely. Camera AI runs sustained workloads — scene detection, multi-frame stacking, HDR — and on an NPU those draw roughly a quarter of the power they would on a GPU, with far less heat.

What's the difference between Apple's Neural Engine, Qualcomm's Hexagon, and Google Tensor's TPU?

Branding and software ecosystem, mostly. All three are dedicated matrix-math blocks inside the main chip. The real differences are which AI frameworks can access them and which features each vendor chooses to gate behind them.

Should I buy a phone with a bigger NPU if I don't use AI features?

Don't pay a premium for TOPS alone. Do weigh NPU generation if you keep phones for three or more years, because new on-device features increasingly require the latest silicon.

Why did my older flagship get excluded from the newest on-device AI features?

Apple and Google now set hard NPU and memory floors for their best on-device models — iPhone 17 Pro/Air for Apple's top tier, Gemini Nano v3 plus 12GB RAM for Google's Gemini Intelligence. Older NPUs fall below the floor, fairly or not.

Does my phone have an NPU, or is that just marketing?

Any flagship from 2023 onward has one. Check your chip's spec sheet for Neural Engine, Hexagon, APU, TPU, or a TOPS rating. On budget phones, "AI" marketing often means cloud processing instead.

Sources

Updated August 2026.

Transparency note: This article was researched and written by Chester Takau with AI assistance for research gathering and drafting. All recommendations reflect the author's own editorial judgment.