NPU, CPU, GPU — and why your IT does not need a garage full of cars
Every new chip generation puts the same question to IT in new clothes: do we now buy everyone the new car? This time the plain-talk answer does not come from the data centre but from the car park — with a garage, a trailer and the question of why car sharing describes a better operating model for the digital workplace than any procurement catalogue.
Three compute engines, three types of vehicle
Before the garage door opens, a quick word on vehicle types. Up to three very different „drivetrains“ work inside every modern computer — and whoever confuses them ends up buying the wrong car.
The CPU is the everyday saloon. Few but strong cores that can drive anything: operating system, applications, logic, databases. Comfortable, universal, built for a single driver — but nobody moves a whole haulage firm with it. Typical models: Intel Core Ultra, AMD Ryzen, Qualcomm Snapdragon X on the client; Xeon, EPYC and Azure Cobalt in the data centre.
The GPU is the full car-transporter convoy. Thousands of simple cores that all drive the same route at the same time. This exact pattern — the same calculation massively in parallel — is why large language models such as GPT, Claude or the models in Azure AI Foundry are trained and operated on GPU clusters (NVIDIA H100/B200, AMD Instinct) in the data centre. A convoy belongs on the motorway, not in a home garage: loud, thirsty, expensive — and irreplaceable when volume has to be moved.
The NPU is the highly specialised efficiency vehicle. A hard-wired compute unit built solely for neural networks, measured in TOPS (trillions of operations per second). It cannot drive an operating system and cannot render a 3D scene — what it does is handle small AI models at a fraction of the energy. Microsoft’s Copilot+ standard requires more than 40 TOPS. The current generation sits well above it: Snapdragon X2 Elite at 80, AMD Ryzen AI 400 at up to 60, Intel „Panther Lake“ at up to 50 — the previous generation at 45 to 50 clears it too. These are vendor figures counted in different ways; what matters is the class, not the decimal place. On top of that run local features such as Windows Recall, Click to Do, live translation and small language models like Phi Silica — not the large models from the cloud.
| Compute engine | In the car picture | Used for | Typical models |
|---|---|---|---|
| CPU | Everyday saloon — can do anything, one lane | OS, apps, logic, orchestration | Core Ultra · Ryzen · Snapdragon X · Xeon · EPYC |
| GPU | Car-transporter convoy — massively parallel | LLM training & inference, rendering, simulation | NVIDIA H100/B200 · AMD Instinct · Azure ND/NC |
| NPU | Efficiency specialist — one job, done perfectly | Local AI: SLMs, Recall, translation, effects | Copilot+ (40+ TOPS): Snapdragon X2 · Panther Lake · Ryzen AI 400 |
The rule of thumb for everything that follows: NPU and cloud GPU are not competitors but different classes of vehicle. The NPU drives small models locally and frugally, the GPU convoy pulls the frontier models in the data centre. The real question is never „which chip is better“ but: which journey is coming up — and which vehicle do you get for it?
The garage, the trailer — and the thinking error behind it
Picture someone with enough money to put the perfect car in the garage for every purpose. For a spirited drive, the AMG C63 or an M3/M4. For off-road, the Toyota Land Cruiser. For the race track, the Ferrari. For the city, the small, nimble city car. Every trip starts by reaching for the right key on the board — as long as only one kind of trip is ever coming up.
The pictureThe garage — and the alternative beside it
On the left is what you own: a separate vehicle for every purpose. On the right is what you can call up when you need it.
But what if the day looks different? Off-road in the morning. Into the city afterwards. To the Nürburgring with friends in the evening. Anyone who has ever argued with a mechanic knows the old-school answer: „Just hitch up the trailer and take both cars.“ Charming — and exactly the moment the model tips over: the effort of hauling the right hardware to the right trip eats up the benefit of specialising.
Each of these machines is the right one for its purpose. The problem is not the individual device — it is the operating model behind it: for every new kind of requirement a new vehicle is bought, per person, for years, sized for the rarest peak. And when requirements change mid-life cycle — and with AI they currently change every quarter — the trailer is wheeled out again: second device, third device, special procurement.
Traffic solved this problem differently long ago: car sharing. Your own everyday vehicle for the daily route — and for the exception the right vehicle is rented, precisely for the hours you need it. Nobody buys a van because they move house twice a year.
The TOGAF question: does everyone now need the NPU car?
With Copilot+, the temptation is back on the table: the new capability sits in the device — so let us swap the fleet. And to be honest: we are making this trip for the third time. Moving from XP to Windows 7, the new system simply demanded more engine — more memory of the new DDR generation, 64-bit platforms; whoever did not upgrade stood still. With the jump to Windows 10/11 came the roadworthiness test: TPM 2.0 and Secure Boot as hard conditions — working devices lost their certificate because a chip was missing. And today the Copilot key sits on the keyboard as the visible sign of the third wave — behind it the NPU as the new entry criterion. Twice the pressure was barely negotiable. The third time, it is fair to ask whether the whole fleet really has to go into the workshop again.
TOGAF has an uncomfortably simple rule for this: in the ADM cycle, technology architecture (phase D) comes after business and information systems architecture (phases B and C). First the capability, then the application, then the technology. Starting with the hardware means buying an answer before the question has been formulated — or in the car picture: ordering the Ferrari before knowing whether anyone wants to go to the race track at all.
- Capability instead of catalogue: „employees can work with AI support“ is the capability being asked for. Whether it is met by a local NPU, a Cloud PC or a browser front end is a solution building block decision — not a procurement reflex.
- Drive the gap analysis honestly: which personas really use NPU features — local inference, offline AI, on-device data protection? For many roles the actual AI load sits in M365 Copilot, Claude or Azure AI: in the cloud, then, not in the chip under the keyboard.
- Decouple the life cycles: devices live four to five years, AI requirements change quarterly. Bolting the two together builds in the next investment backlog straight away.
That does not make NPUs unnecessary — they belong where the analysis puts them: with personas that genuinely need local inference, offline capability or data processing that must not leave the device. For the breadth of the organisation there is the car-sharing model — and it is becoming very concrete right now.
When agents ride along: the new load profile
Why is the everyday car suddenly not enough? Because the journeys have changed. In 2026, AI increasingly works in parallel with the person — as an agent that works through tasks on its own while the person has long moved on to something else. With Claude Cowork, Anthropic’s desktop app for agentic knowledge work, users delegate whole packages of work — research, document analysis, file handling — to an agent in the background. ChatGPT offers the same pattern with its agent mode, Microsoft Copilot brings agents into the M365 stack via Copilot Studio and Agent 365, and on Azure AI Foundry organisations build their own agents for their business processes.
In the car picture: it is no longer only the person driving. Alongside your own trip, one or two further vehicles are constantly out on errands — and they need sometimes a small car (tidying up a file store), sometimes the van (analysing 500 pages), sometimes the convoy (processing images and video). For IT that means, concretely:
- Parallelism: one employee generates two to three sessions‘ worth of load — their own work plus running agent tasks.
- Volatility: agent workloads come in bursts. An hour at full load, then hours of idling — the opposite of predictable commuter traffic.
- Shifting kinds of requirement: today storage and I/O, tomorrow compute, the day after a GPU. The performance requirement is no longer a constant but a curve.
- Governance: agents need identity, policies and isolation just like human users — Entra ID, Intune and Conditional Access become the traffic code for software.
A vehicle bought outright is sized for exactly one load profile. This working world no longer has one.
Car sharing instead of a trailer: rent what the journey needs
That sets the strategic points. There are two answers to the new AI load:
Path A — extend the garage: buy the right device for every kind of requirement, per person, and run the next fleet refresh with every new class of requirement. That is the trailer model: technically feasible, economically and organisationally always one cycle too late.
Path B — the car-sharing model: every person gets one fixed vehicle for the main working hours — either as a solid hardware device or directly as a permanently assigned virtual machine (Cloud PC or AVD desktop). For everything that deviates from that, nothing is bought — it is rented for the time it is needed: a GPU-accelerated session for Tuesday’s 3D simulation, an AI-enabled Cloud PC for the new Copilot scenario, an isolated agent environment for the month-end close. Alternatively — for external staff on personal devices, for instance — two virtual machines: one for daily work, one for agents and special loads.
The decisive mechanism behind it: with a virtual machine, the hardware underneath the workplace can be changed without touching the workplace. From 4 to 8 vCPU, from standard to GPU sizing, from a classic to an AI-enabled Cloud PC — by configuration in minutes instead of by a procurement project in months. If a new AI requirement arrives tomorrow, you rebook rather than buy. That is precisely the difference between a fleet you own and mobility you subscribe to.
And yes — there are still good reasons for your own specialist vehicle: development with local inference, field staff without reliable connectivity, privacy scenarios with on-device processing. The car-sharing model does not ban the Ferrari. It only prevents the whole workforce from getting one because three people go to the race track.
From the car park to the nautical chart
If you know the nautical chart metaphor from this series, you can draw the car-sharing picture straight onto it. Two currents currently pull in opposite directions: the coastal current „on-device“ — chip makers and Microsoft are pushing AI out across the device estate, promising near-zero latency and data that never leaves the machine. And the open-sea current „cloud & agents“ — the large models, agent platforms and elastic compute grow in the data centre faster than any device fleet could be retrofitted.
A ship that tries to run both currents with a single, heavily laden hull will be fast in neither — that is the garage at sea. The seamanlike answer is the same as in the car park: a nimble boat for the coast, access to the big ship for the open sea. The shallows are marked: hardware bets on a single model generation, lock-in to one load profile, fleet refreshes whose justification is already out of date at rollout. The lighthouse stays the same as in the nautical chart article: build decisions so that they stay reversible.
Where this is heading: virtual desktops in the age of agents
Perhaps the most important development of the year: the virtual desktop is itself changing role — from a workplace for people to an execution environment for people and agents. Four markers show the direction:
1 · AI-enabled Cloud PCs: Windows 365 brings Copilot+-style features — improved Windows search, Click to Do — into the Cloud PC. The „NPU“ then no longer sits in the device on the desk but in the Azure infrastructure underneath. In front of it can stand a five-year-old notebook or a Windows 365 Link — the AI capability still arrives. That is the car-sharing idea in its purest form: the vehicle class changes without anyone buying a new car.
2 · Windows 365 for Agents: generally available since Build 2026, AI agents receive their own Cloud PCs within Agent 365 — Microsoft’s control plane for enterprise agents — full, isolated Windows or Linux environments in which they operate applications, process files and run multi-step workflows, including in legacy systems without APIs. Governed with the same tools as human users: Entra ID, Intune, policies. The Cloud PC becomes the rental fleet for software.
3 · Agents inside your own Cloud PC: assistants such as Microsoft Scout — Microsoft’s first „autopilot“ agent, as of September 2026 an experimental preview in the Frontier programme and not yet generally available — show the direction: commissioned by Teams chat from a phone, the agent carries on working in the personal Cloud PC or desktop — find a file, summarise it, prepare an email draft. The virtual desktop keeps driving even when nobody is at the wheel — and that is precisely what a device cannot do by design.
4 · The direction behind it: Microsoft now frames the strategy across platforms itself — AI workloads should be able to run on the device, in the cloud or across both, without committing in advance. That is exactly why the placement question matters more than the chip question: the future belongs not to one location but to the ability to choose the location per workload — and to change it. Anyone building architecture today builds the switch in from the start.
The same logic applies to AVD with a greater depth of control: assign GPU sizings specifically to the personas that need them, isolate agent workloads in their own host pools, let capacity follow the actual curve through autoscaling — instead of imposing the most expensive common denominator on the whole fleet.
The conclusion — no detours
Buy the vehicle for the main route. Rent the rest.
First: CPU, GPU and NPU are vehicle classes, not competitors — saloon, convoy, efficiency specialist. The question is always the journey, never the chip.
Second: by TOGAF logic, a blanket NPU rollout is a phase D decision — it is taken after the capability and gap analysis, not in the procurement catalogue.
Third: agentic work with Claude Cowork, ChatGPT, Copilot and Azure AI makes load profiles parallel, bursty and changeable — the opposite of what you buy hardware outright for.
Fourth: the car-sharing model is the most robust answer: one fixed vehicle for the main hours — hardware or virtual — and for everything else rent the capacity you need for as long as you need it. If the requirement changes, the hardware under the virtual machine is rebooked instead of the garage being extended.
What the two operating models really cost across the life cycle belongs in an honest TCO analysis — from device price through utilisation to refresh cycles. And if you want to see which of the six DaaS platforms covers which sizing, sharing and agent scenario today: the DaaS Maps set out more than 270 features vendor-neutrally.