Choose the right power architecture for your NPU-equipped MCU. Compare on-chip regulators, external PMICs, and hybrid supplies for inference burst workloads.
Your always-on sensor node now comes with a neural accelerator, and the first prototype browns out mid-inference. The battery still had charge, but the rail could not hold through the inference burst. As covered in Your MCU Got Smarter – Your Supply Chain Got Harder, some of the latest NPU-equipped MCUs now cost only a dollar or two, bringing sharp, millisecond-scale inference bursts into designs that may have been sized for relatively steady control workloads.
A neural accelerator on a tight power budget is a power problem before it is a compute one. You have three basic choices: use the MCU's on-chip regulators, move the core rails to an external power-management IC (PMIC), or combine internal and external supplies in a hybrid architecture. Each involves different trade-offs in component count, board area, efficiency, rail control, and performance headroom.
An always-on device usually keeps a low-power sensing path awake while the main compute domain waits for a trigger. When the trigger fires, inference runs in a burst. Chipmaker Alif's event-level breakdown of a keyword-spotting inference cycle separates the workload into a wake spike, the mel-frequency cepstral coefficient (MFCC) processing burst, the NPU burst, and the tail current that follows. Alif measured the workload at 6.8 ms and 0.17 mJ per inference on the NPU, compared with 137 ms and 2.62 mJ on the CPU alone.
A 2025 independent preprint measuring Alif silicon with external instrumentation found per-inference energy improvements ranging from 8.9x to 143x across four models that benefited from NPU offload.
In a keyword-spotting example, average power during the NPU run was somewhat higher, but the much shorter execution time more than compensated.
On a tiny MNIST handwritten-digit classification model, the NPU ran at just 5.1 percent utilization and consumed more energy than the CPU path because orchestration overhead outweighed the compute savings. For small models, the CPU can be more efficient; where the NPU has independent power control, firmware can leave it off until the workload warrants acceleration.
Measured NPU Energy Results Across Workloads
Example | Measured Result | Design Lesson |
Alif keyword spotting | CPU: 137 ms, 2.62 mJ | NPU finishes about 20x faster and uses about 1/15th the energy. |
2025 independent preprint | 8.9x to 143x energy improvement across four NPU-beneficial models | Efficiency depends strongly on model fit and accelerator utilization. |
Tiny MNIST digit classifier | NPU utilization: 5.1%; energy use exceeded the CPU path | Small models can lose to orchestration overhead. |
NPU efficiency depends heavily on workload fit and accelerator utilization.
To see what this means at the power rail, look at ST's STM32N6. It runs the core at 0.81 V in nominal mode, and ST's measurements put the VDDCORE domain alone at 124 to 191 mW during inference. The company ships a dedicated power-measurement package for this rail.
Public MCU-NPU power data rarely include oscilloscope-level current slew rates or rail-droop traces. Fast load steps are a known problem, but the transient your regulator and decoupling must handle depends on the workload, board, converter, and source. Measure VDDCORE on the actual hardware under inference.
Once you know what the load looks like, the question becomes where the conversion happens: on the chip, in an external PMIC, or split between the two.
MCU vendors have been putting more of the power path onto the chip for years, and the NPU generation continues that trend. ST's STM32U5 and STM32N6 families offer internal switching regulation for their core supplies, requiring only a small set of external inductive and decoupling components. Infineon's PSOC Edge splits the chip into a high-performance domain and a low-power domain.
Looking ahead to 2027, Ambiq says its upcoming Atomiq110 will use a 12 nm implementation of its Subthreshold Power Optimized Technology (SPOT) platform and feature an ultra-low-power mode operating near 300 mV.
NXP's i.MX RT700 makes the trade-off unusually visible because its evaluation kit (EVK) lets designers evaluate the same processor under four power architectures: a fully external PMIC, a fully on-chip design, a hybrid of the two, and a discrete build with no PMIC. On external rails at 1.1 V or higher, the main core reaches 325 MHz and the sense core 250 MHz. Using the internal low-dropout regulators (LDOs) lowers those ceilings to 250 and 205 MHz. The external-PMIC configuration ships as the EVK's default. Rail count is therefore only part of the make-or-buy decision, because an external supply also buys measurable clock headroom and can improve conversion efficiency.
On-chip regulation is limited by current and thermal headroom, conversion efficiency on the system-on-chip (SoC) process itself, and the number of rails it can control. Still, the internal path minimizes external components and board area, and for many sub-watt designs, it is the better fit, provided the on-chip regulator can support the required current, rail stability, and clock targets.
For endpoint designs, an external PMIC can improve conversion efficiency and add rail control, current capability, sequencing, and functions that the MCU does not provide internally. In rechargeable systems, it can also integrate battery charging and gauging.
Nordic's nPM1300 shows how much power-management functionality can fit into a small package. It combines two 200 mA buck regulators, two configurable LDO/load switches, battery charging, and fuel-gauge support in a package around 3 by 2.4 mm. It is configured over I2C and can operate with as few as five external passives.
A dedicated power die can regulate more efficiently than circuitry built on the MCU's logic process. Those gains must justify the added board area, components, and cost.
A hybrid architecture keeps some conversion on the MCU while moving the rails that benefit most from external regulation off-chip. NXP's i.MX RT700 EVK provides a useful example: its on-chip DC-DC converter generates VDDN while external regulators supply VDD1 and VDD2. A hybrid configuration fits when one or two rails need the efficiency or performance headroom of external regulation, while the remaining rails do not justify the board area and cost of a fully external PMIC architecture.
The cleanest BOM comparison uses the same processor under different power configurations, and NXP's i.MX RT700 provides that. The external path adds the PCA9422 and its passives to the BOM in exchange for greater clock headroom and potentially higher conversion efficiency. The on-chip path removes those line items but gives up some performance headroom. List and spot prices have sometimes diverged sharply this year, so use Octopart to price both power stacks at the quantities you expect to purchase.
For the second-source and qualification side of MCU selection, see 5 Suppliers Hold 80% of the MCU Market. Where's Your Second Source?
A coin cell's internal resistance can rise from several ohms to tens of ohms as it discharges, so an inference burst drawn directly from the cell can drop the converter input below its required operating voltage. A reservoir capacitor or small supercapacitor can supply the burst locally. Size it from the burst's charge demand, battery internal resistance, and allowable voltage droop rather than from peak current alone.
TI's long-standing white paper SWRA349 walks through the calculation. Validate the reservoir capacitor at the battery's worst-case state of charge and operating temperature, when source impedance is highest. In a marginal design, this one component can determine whether the product ships.
When it comes time to choose a power architecture, start with the power class. A sub-watt design optimized for minimum board area and BOM cost usually starts with the on-chip regulator. A multi-domain or performance-sensitive SoC is more likely to justify an external PMIC or hybrid supply through improved efficiency, sequencing, rail control, or clock headroom. High-impedance sources such as coin cells need burst buffering regardless of the regulator architecture.
Whichever architecture you choose, two measurements should drive the decision. Energy per inference shows whether acceleration improves the energy budget. The worst-case rail transient shows whether your regulator, source, and decoupling can support the workload. Measure both on your own hardware before you commit.