Back to News
User submission36kr.com

AMD Pushes Personal AI With Local Models and “Token Freedom”

Summary

AMD is positioning local computing as the foundation of a “personal AI” era, arguing that more intelligence should run on PCs, laptops and workstations instead of only in data centers. Its new Gorgon Halo platform raises unified memory from the 128GB level associated with Ryzen AI Halo and supports local deployment of models with up to 300 billion parameters, while AMD presents it as a response to Nvidia’s 128GB DGX Spark and its cited 200-billion-parameter ceiling. AMD executive Jack Huynh frames personal AI around three elements: local computing, privacy and user control, and personal context. The company says monthly token processing could rise from about 0.7 quadrillion to 1.7 quadrillion and reach 120 quadrillion by 2030, making the cost and location of inference more important. AMD cites software-engineering comparisons in which locally run Laguna S 2.1 on Strix Halo outperformed cloud-based Claude Sonnet 5, and GLM 5.3 Flash on Gorgon Halo outperformed Claude Fable 5; these claims are presented in the article without detailed test methodology. The company says Gorgon Halo can support devices such as Lenovo ThinkCentre X and HP’s ZBook, while Gorgon Halo Box is intended for developers and is planned for Europe. Microsoft’s Project Zenith is described as a preconfigured Windows programming environment with VS Code, WSL, GitHub Copilot CLI and PowerShell, designed to support local large-model and complex AI workflows on Ryzen AI Halo systems. At the high end, Threadripper Halo Station combines a 96-core Threadripper PRO processor, two expandable Instinct MI350P accelerators, up to 2TB of system memory and 576GB of HBM3E to support trillion-parameter models locally. AMD’s broader proposition is a graduated path from a personal AI PC to a workstation and then larger infrastructure, with workloads split between local devices and data centers according to their needs. The article also emphasizes that local deployment can improve privacy and avoid per-token billing, but the cited hardware cost comparison shows that local inference is not automatically cheaper in every usage scenario.