AMD Puts 96 Cores and 576GB of HBM3e Beside Your Desk at IFA: Trillion-Parameter Models Locally, Says AMD — but the Demo Unit Held Two Cards

1 views

At the IFA 2026 opening keynote on September 4, AMD's Jack Huynh introduced the Threadripper Halo Station, calling it the most powerful workstation on the market and saying it can run AI models with more than a trillion parameters locally. The configuration is a 96-core, 192-thread Ryzen Threadripper PRO 9995WX with up to 2TB of eight-channel DDR5, plus PCIe-form Instinct MI350P accelerators — 144GB of HBM3e each, 4 TB/s of bandwidth, and up to 4.6 petaFLOPS of FP4 compute per card. Four cards make 576GB of HBM3e and a combined 16 TB/s, enough to hold a trillion-parameter model entirely in GPU memory at four-bit precision. But the unit shown at IFA had room for only two cards, both liquid-cooled, and AMD says only that there is "a path" to installing four. No price or release date has been announced; press estimates run $100,000–$150,000, shipping next year.

Three Conditions Attached to "Trillion Parameters"

The numbers themselves hold up. The MI350P has 128 CDNA 4 compute units on TSMC N3 and is essentially half an MI350X in PCIe form, with 144GB of HBM3e per card. Four of them give 576GB and 16 TB/s, and a trillion-parameter model quantized to four bits does fit in that. But the claim carries three conditions. Precision has to be four-bit, not FP16. All four slots have to be filled — and the IFA demo unit had room for two, with AMD's wording being that there is "a path" to four rather than "we sell a four-card version." And there is power: four 600W cards push past what a standard North American outlet supports unless the cards drop to 300W or the install mandates a 20-amp circuit. The demo holding two cards is probably not a coincidence. MI350P also doesn't sell through consumer channels; outside estimates put a single card near $20,000, so the system price can only be inferred. AMD's comparison target is Nvidia's DGX Station — a 252GB B300, a 72-core Grace CPU and 496GB of LPDDR5x for around $100,000 — against which AMD claims up to 3.4× the system memory and more than twice the memory bandwidth.

There Is Nothing New in It, and That Is the Point

Taken apart, this machine recombines parts that already exist: an off-the-shelf Threadripper PRO, off-the-shelf MI350 silicon, and a liquid-cooling loop. The variable is not the hardware but the price point at which both vendors are now pushing rack-class capacity next to a desk — six-figure desk-side machines are becoming a real product category rather than a special order. Note how differently it lands from the other two local-AI stories of the same week. This site covered NVIDIA's PAIR yesterday, which pools idle household machines at zero cost but does not raise the ceiling on any single request, and Anker's 26 TOPS home hub from the same show the day before. Halo Station sits at the far end of that line: what it buys is not cheapness, but fitting a model too large for ordinary hardware entirely into local memory.

Who Actually Needs One

Start with who doesn't: teams trying to save on API bills. At $100,000–$150,000 plus power and maintenance, that money buys far more inference from a cloud endpoint in almost every scenario. The fit is narrower — data that genuinely cannot leave the building (medical, legal, classified engineering), long jobs run repeatedly without queueing for shared capacity, or local fine-tuning and layer-by-layer debugging of large models. What those share is that the purchase is about determinism and control, not compute per dollar. Two things still have to land: there is no official price or ship date, and no clear word on whether a four-card configuration will actually be sold. Until those settle, it is early to write this into a 2027 procurement plan.

via: The Register, Tom's Hardware, TechPowerUp