I Tried to Run a Language Model on My Laptop’s NPU. It Fell Back to the CPU 96% of the Time.
Laptop NPU test: Every AI laptop sold in the last two years has a number on the box: 45 TOPS. That’s the neural processing unit — the NPU — the dedicated silicon that’s supposed to make on-device AI fast and efficient. It’s the reason the marketing says “AI PC.”
I have one. A Surface running a Snapdragon X Elite: 12 CPU cores, an Adreno GPU, and that 45-TOPS Hexagon NPU, all sharing 64 GB of memory. I wanted to run a language model on it — locally, no cloud, no API bill — and I wanted to use the NPU, because that’s the part everyone’s paying for.
So I spent a few weeks trying. This is the autopsy. It’s an honest negative result, which is the most useful kind, because nobody publishes them and everybody needs them.
The short version: on the standard software path, 96.2% of the model ran on the CPU anyway, and the part that reached the NPU couldn’t generate text at all. The longer version is more interesting, because the reason why tells you something true about these machines that the spec sheet actively hides.
The promise, and the first crack
The pitch for an NPU is simple and, in the abstract, correct: matrix multiplication is what language models do, NPUs are built for matrix multiplication, and they do it at a fraction of the energy a CPU spends. On paper, the NPU is the obvious place to run a model.
To get a model onto the Hexagon NPU, the supported route is Microsoft’s ONNX Runtime with Qualcomm’s execution provider — the software layer that’s supposed to hand your model’s math to the NPU. I took a standard, off-the-shelf quantized model (a small Qwen 2.5, the kind you’d download in thirty seconds) and pointed it at that path.
Here’s what the runtime actually did with it. The model, once compiled, was a graph of 2,918 operations. The NPU accepted 112 of them. The other 2,806 — 96.2% — silently fell back to the CPU. And the mixed path wasn’t just mostly-CPU; it was slower than just using the CPU directly: 1,105 milliseconds to process a prompt versus 172 milliseconds CPU-only. More than six times slower, to use the accelerator.
Then I tried to actually generate text — the token-by-token loop that is the entire point of a language model. It failed outright, with a shape-mismatch error deep in the attention math. The NPU could do a single forward pass on a toy input. It could not run a real model’s generation loop.
That’s the headline. But “the NPU doesn’t work” is the wrong lesson, and if I’d stopped there I’d have published something false.
Why it fell back: the model wasn’t shaped for the silicon
The NPU didn’t reject the math because it couldn’t do the math. It rejected the shape of the graph.
An off-the-shelf downloaded model is built for flexibility. It uses dynamic shapes — the sequence length can be anything, the batch can be anything, the growing memory of past tokens can be anything — and it leans on a Microsoft-specific way of packing its compressed weights. That flexibility is exactly what makes it painless to run on a CPU or GPU.
The NPU wants the opposite. It wants a graph frozen into fixed shapes, quantized in a specific format, and compiled ahead of time into a binary built for that exact chip. Hand it a flexible, dynamic-shape model and it can only take the handful of operations that happen to be static — 112 of them — and it shrugs the rest back to the CPU.
So the failure was real, but its cause was model format, not silicon capability. The NPU can run a language model. It cannot run this one, downloaded this way, through this runtime. (Later work confirmed the other direction: models specifically compiled through Qualcomm’s own toolchain into fixed-shape binaries do run on the NPU — slowly, but they run. The generic path is the trap, not the chip.)
That distinction matters, because it’s the difference between the hardware is a lie and “the ecosystem hasn’t connected the pieces yet.” The truth is the second one. But there’s a deeper reason the NPU wouldn’t have saved me even if the format had matched — and that one is physics.
The part the spec sheet hides: one bus, shared by everyone
Here’s the architectural fact that reframes the whole exercise. On this chip, the CPU, the GPU, and the NPU don’t have their own private memory. They share a single pool — 64 GB — over a single memory bus rated at 135 GB/s.
Generating text with a language model is not, mostly, a compute problem. It’s a memory problem. To produce each token, the machine has to read the model’s entire multi-gigabyte set of weights out of memory. The math on those weights is trivial by comparison. So the thing that limits your speed isn’t how many TOPS your accelerator has — it’s how fast you can stream weights across that shared bus.
And that reframes everything. The NPU’s 45 TOPS advantage is an advantage at the thing that wasn’t the bottleneck. When three compute units all read from the same bus, adding the NPU to the CPU’s job doesn’t give you a second lane — it gives you two tenants fighting over the one lane you already had.
I tested this directly. I wanted to know: if I run a “zero-copy” design — allocate memory once, let every compute unit read it in place, never copy anything — how much does that buy? The answer, measured, was startling in the wrong direction.
– The APIs that would let the CPU and GPU coherently share a pointer into the same memory? Absent on this platform. The shared-memory capability flag reads back as zero. The one path that looked like it worked failed a correctness check at every size I tried.
– The API to register memory directly into the NPU without a copy? Not present in the shipped runtime at all. You need Qualcomm’s full professional SDK, which isn’t what the machine comes with.
– And the one true zero-copy path that did work — sharing a buffer between two graphics APIs — was slower than just copying the data. At the size of a single token’s worth of data, copying took 102 microseconds. “Zero-copy” sharing took 212, because the cost of synchronizing the handoff between the two APIs was bigger than the copy it was supposed to avoid.
That last one is the finding I’d frame and hang on the wall: zero-copy is not a virtue. It’s a measurement. Avoiding a copy is only faster if the copy was the expensive part. On this hardware, at these sizes, it wasn’t.
What actually worked (it’s boring)
After weeks of chasing the accelerators, the fastest, most reliable way to run a language model on this “AI PC” turned out to be: the CPU. Just the twelve Oryon cores, running a well-optimized open-source runtime, at four threads.
I benchmarked every model I had across CPU and GPU. The CPU won token generation on every single one — often by a lot. A 4-billion-parameter model generated at 28.7 tokens/second on the CPU versus 16.2 on the GPU. A 26-billion-parameter model: 21.3 on the CPU, 8.4 on the GPU. The CPU’s cache hierarchy quietly absorbs the weight reads that the GPU has to fetch across that shared bus, and it wins.
There’s one genuinely useful thing the accelerators buy you, and it’s not speed — it’s the ability to run *two* models at once. If you put your main model on the GPU and a small background model on the NPU, the GPU foreground degrades less than if you’d paired that background model with the CPU. It’s not free — the foreground takes a real hit — but for running a big writer and a small helper side by side, splitting them across lanes is the right instinct. That’s the actual, defensible version of “use the whole chip”: not one model spread across three units, but different models in different lanes, chosen by which lane survives the contention.
The lesson worth keeping
There are two different gaps in this story, and it matters which is which.
The first is why the NPU wouldn’t run my model — the format mismatch, the 96.2% fallback. That gap is software. It’s the ecosystem not having connected the pieces yet, and it will close: compile a model correctly for the chip and the NPU does run it. Software you can fix.
The second gap is why nothing was fast in the first place — and that one is hardware, and you can’t fix it. Every part of this chip reads from a single memory bus rated at 135 GB/s. Generating text is bottlenecked by how fast you can stream a model’s weights across that bus, not by how much compute you can throw at them. So the ceiling on how fast any model runs on this machine isn’t set by the CPU, the GPU, or the NPU. It’s set by that one number — and no amount of clever software moves it. Better quantization, sharper flags, mixture-of-experts models, unified runtimes: all of that gets you closer to the ceiling. None of it raises the ceiling.
Here’s the proof. Take the exact same model, the exact same software, and move it to a Mac mini M4 Pro — a machine whose unified memory runs at roughly 273 GB/s, about double this Snapdragon’s bandwidth. You’d expect roughly double the token speed, and the felt experience lands on the other side of a line: sluggish becomes fluid. Nothing changed but the bus. That’s the whole point. The number that governs your experience is memory bandwidth, and it’s a property of the silicon you bought, not the code you write.
Which means the marketing puts the wrong number on the box. It says 45 TOPS — a compute figure, at the thing that was never the bottleneck. The number that would actually tell you how a language model feels on this machine is GB/s, and you’ll never find it on the sticker.
If you own one of these machines and you want to run models locally: use the CPU, use a good runtime, and treat the NPU as a specialist you’ll grow into, not the engine you start with. Software will get you to the edge of what the hardware allows — and then the memory bus decides the rest.
And if a benchmark ever tells you the accelerator is faster, check what fell back to the CPU while you weren’t looking.
Mine did. 96.2% of it.
—
I run a small multi-agent research setup at home and publish what I find, including — especially — the things that didn’t work. All the numbers here trace to on-device measurement logs; if you want the raw benchmark data or the reproduction steps, get in touch.
Leave a Reply
Want to join the discussion?Feel free to contribute!