AMD Unveils Threadripper Halo Station: A Trillion-Parameter AI Workstation
AMD's new liquid-cooled Threadripper Halo Station, featuring a 96-core Zen 5 CPU and dual MI350P accelerators, is hailed as the world's most powerful workstation, capable of running trillion-parameter AI models locally.
✨ This content was summarized and interpreted by AI; it may contain errors — please verify accuracy with the original sources. Learn more
Listen to this story

The Threadripper Halo Station, unveiled by AMD at IFA 2026, directly targets the burgeoning market for high-performance AI workstations with an unprecedented combination of processing power and memory, positioning itself as "the most powerful workstation in the world" capable of running trillion-parameter models. This liquid-cooled powerhouse features a 96-core Zen 5 Threadripper PRO 9995WX processor, two liquid-cooled AMD Instinct MI350P accelerators with a planned upgrade path to four, a staggering 2TB of DDR5 system memory, and 288GB of HBM3E memory from the dual GPUs, expandable to 576GB with four accelerators.
The core of this system's immense capability lies in its components. The Ryzen Threadripper PRO 9995WX, a Zen 5 chip, boasts 96 cores and 192 threads, with boost clocks up to 5.4 GHz and 384 MB of L3 cache, operating within a 350W TDP. Zen 5 architecture brings significant refinements, including improved branch prediction, wider instruction dispatch, and a fully optimized AVX-512 implementation, enhancing instruction throughput and floating-point intensive workloads crucial for AI. Each MI350P accelerator, based on AMD's CDNA4 architecture and built on TSMC's advanced 3nm and 6nm FinFET processes, comes with 144GB of HBM3E memory, delivering 4 TB/s bandwidth and a 128MB last-level cache. The MI350P offers up to 4.6 PFLOPS of MXFP4 or MXFP6 matrix performance and 2.3 PFLOPS in MXFP8, making it roughly 40% faster in FP16 and FP8 theoretical compute performance compared to Nvidia's H200 NVL. This performance is particularly crucial for inference and Retrieval-Augmented Generation (RAG) pipelines, which are key to many modern AI applications.
This unveiling marks a significant escalation in the workstation market, demonstrating AMD's aggressive push into local AI. The ability to run trillion-parameter models locally addresses a critical need for AI researchers, data scientists, and developers who require immediate, private, and unmetered access to powerful compute for model training, fine-tuning, and inference without the latency and recurring costs associated with cloud-based solutions. AMD's emphasis on "Personal AI" with local compute, privacy, and personal context resonates with a growing sentiment that not all intelligence needs to be rented from large cloud providers. The market for AI is projected to reach $30 trillion, and while cloud AI remains dominant for hyperscale training, the demand for powerful local AI solutions is rapidly expanding, especially as agentic AI drives token consumption to unprecedented levels.
In comparison to rivals, the Halo Station directly challenges Nvidia's stronghold in the high-end AI workstation segment. Nvidia's DGX Station for Windows, powered by the GB300 Grace Blackwell Ultra Desktop Superchip, also claims trillion-parameter model support with 748 GB of coherent memory and up to 20 petaFLOPS of AI compute performance. While Nvidia's solutions leverage a highly mature CUDA ecosystem, AMD's MI350P, with its CDNA4 architecture, has shown impressive raw performance gains, particularly in FP8 and FP16, and its ROCm software stack has seen significant improvements, now supporting frameworks like PyTorch and JAX with reduced configuration effort. The MI350P's 144GB HBM3E memory per card provides a substantial capacity advantage over Nvidia's H200 (141GB) for fitting larger models on a single GPU, which can translate to a lower hardware cost per token for certain large models. However, the MI350P currently lacks the Infinity Fabric GPU-to-GPU interconnect found in its OAM counterparts (MI350X/MI355X), limiting multi-GPU communication to PCIe Gen5 x16, which could impact usable AI model sizes in multi-GPU setups. Nvidia's B200, while having less memory per card (180GB), offers 8.0 TB/s of bandwidth and NVLink 5 at 1.8 TB/s for multi-GPU scaling. Intel is also making strides in the AI PC and workstation space, showcasing Core Ultra Series 3 processors with integrated AI acceleration and new Intel Arc Pro B70 and B65 GPUs, alongside Xeon 600 workstation processors.
Looking ahead, the Threadripper Halo Station represents a strategic move by AMD to capture a significant share of the rapidly expanding AI workstation market. The liquid cooling solution, essential for managing the 600W TBP of each MI350P and the 350W TDP of the Threadripper CPU, underscores the extreme performance levels targeted. While AMD has not yet announced pricing, estimates suggest a street price well over $100,000, potentially climbing to over $150,000 with full configuration, placing it in direct competition with Nvidia's high-end DGX solutions. The success of the Halo Station will hinge not just on raw benchmarks, but on the continued maturation of the ROCm software ecosystem and robust OEM partnerships to bring these systems to market effectively. The planned support for up to four MI350P accelerators suggests a future roadmap focused on even greater memory capacity and compute density, essential for the ever-growing demands of AI model development. This push for local, unmetered AI compute on powerful workstations will likely accelerate the decentralization of AI development, empowering more users to innovate outside of traditional cloud infrastructure.