Inference moves onto the device
Small models running locally are about to reprice the unit economics of consumer AI.
The attention stays on frontier models, but the margin story is happening at the other end of the size curve. Models under ten billion parameters now hold most of the capability that consumer applications actually use: summarization, drafting, translation, voice, structured extraction. At the same time, the neural compute shipping in phones and laptops has roughly doubled every two years. Those two curves crossed quietly. A task that cost a server round trip in 2023 runs on the handset today.
The pattern to watch is not model quality. It is who pays for inference. Cloud inference is a marginal cost that scales with usage, which is why most consumer AI products cap usage or lose money on their heaviest users. On-device inference inverts that: the customer bought the compute when they bought the hardware. The marginal cost of one more request is zero to the developer.
Three consequences follow.
First, pricing pressure on API-first products. Any feature that a seven billion parameter model can serve becomes hard to charge for, because a competitor can ship it at zero marginal cost. The defensible surface shrinks to tasks that genuinely need frontier scale, fresh data, or proprietary context.
Second, privacy stops being a compliance line and becomes a product feature with real distribution behind it. Health data, financial records, personal messages: the categories where users hesitate to send data to a server are exactly the categories where local inference removes the objection. Expect the first durable consumer AI brands in those categories to be local-first, and expect the platform owners to market that distinction hard.
Third, the telecom and edge story gets simpler, not bigger. If the handset does the work, the operator does not get a new revenue line from edge inference. The winners of the shift are the chip designers with NPU share and the platform owners who control the local model runtime and take a toll on distribution.
The signal to track over the next four quarters: how aggressively the two dominant mobile platforms restrict or tax third-party access to their on-device models. Openness there decides whether local inference becomes a commodity layer or another walled garden. If access stays cheap and broad, thousands of small products get viable overnight. If it narrows, the platforms capture the shift and the API providers keep the middle of the market.
Either way, businesses paying per token for commodity tasks should treat that line item as temporary. The work is moving to hardware the customer already owns.
Get the next one the moment it drops.
Free account, one email per published piece, unsubscribe in one click.