iPhone 17 Pro Demoed Running a 400B Model

A demo showed an iPhone 17 Pro running a 400B-parameter large model. Behind the gimmick of a phone running a flagship-class model, the technical details are more interesting than the conclusion.

The Quality of the Demo

First, let's put the question marks on the table: a 400B-parameter model at normal precision needs hundreds of GB of memory, and a phone's physical ceiling is what it is, so a demo like this necessarily relies on some combination of extreme quantization, sparsification (the parameters actually activated being far fewer than the total), or partial computation offloading. A demo running through and everyday usability are two different things—the cost in generation speed, heat, and battery is usually cut out of the demo video. Half the community discussion is in awe, and half is pressing on those omitted specs; the latter half is more worth reading.

The Direction Is More Real Than the Numbers

Strip away the marketing fluff, and the trend itself holds up: Apple Silicon's unified memory architecture happens to suit large-model inference, and the memory bandwidth of the A-series and M-series has been raised generation over generation, clearly paving the way for on-device AI; the MoE architecture makes "big parameters, small activation" possible, and the roadmap for on-device flagship models is no longer science fiction. The real significance for users comes in a year or two: truly private AI (data never leaves the device), an assistant usable without a network, and zero-marginal-cost local inference. How much of today's demo is performance doesn't matter; what matters is that on the direction it demonstrates, every chip maker is placing real, hard-cash bets.

via: Hacker News