Three Conditions Attached to "Trillion Parameters"
The numbers themselves hold up. The MI350P has 128 CDNA 4 compute units on TSMC N3 and is essentially half an MI350X in PCIe form, with 144GB of HBM3e per card. Four of them give 576GB and 16 TB/s, and a trillion-parameter model quantized to four bits does fit in that. But the claim carries three conditions. Precision has to be four-bit, not FP16. All four slots have to be filled — and the IFA demo unit had room for two, with AMD's wording being that there is "a path" to four rather than "we sell a four-card version." And there is power: four 600W cards push past what a standard North American outlet supports unless the cards drop to 300W or the install mandates a 20-amp circuit. The demo holding two cards is probably not a coincidence. MI350P also doesn't sell through consumer channels; outside estimates put a single card near $20,000, so the system price can only be inferred. AMD's comparison target is Nvidia's DGX Station — a 252GB B300, a 72-core Grace CPU and 496GB of LPDDR5x for around $100,000 — against which AMD claims up to 3.4× the system memory and more than twice the memory bandwidth.
There Is Nothing New in It, and That Is the Point
Taken apart, this machine recombines parts that already exist: an off-the-shelf Threadripper PRO, off-the-shelf MI350 silicon, and a liquid-cooling loop. The variable is not the hardware but the price point at which both vendors are now pushing rack-class capacity next to a desk — six-figure desk-side machines are becoming a real product category rather than a special order. Note how differently it lands from the other two local-AI stories of the same week. This site covered NVIDIA's PAIR yesterday, which pools idle household machines at zero cost but does not raise the ceiling on any single request, and Anker's 26 TOPS home hub from the same show the day before. Halo Station sits at the far end of that line: what it buys is not cheapness, but fitting a model too large for ordinary hardware entirely into local memory.
Who Actually Needs One
Start with who doesn't: teams trying to save on API bills. At $100,000–$150,000 plus power and maintenance, that money buys far more inference from a cloud endpoint in almost every scenario. The fit is narrower — data that genuinely cannot leave the building (medical, legal, classified engineering), long jobs run repeatedly without queueing for shared capacity, or local fine-tuning and layer-by-layer debugging of large models. What those share is that the purchase is about determinism and control, not compute per dollar. Two things still have to land: there is no official price or ship date, and no clear word on whether a four-card configuration will actually be sold. Until those settle, it is early to write this into a 2027 procurement plan.