Apple Reportedly Plans Baltra AI Inference Server With M8 Ultra Chips
Summary
Apple is reportedly discussing Nvidia’s NVLink Fusion technology for a planned enterprise AI inference server, a move that could mark its return to the server hardware market. The system, internally known as Baltra, is being developed with Broadcom and may be offered in configurations using two or four of Apple’s in-development M8 Ultra chips. The server is intended for AI developers, businesses and government organizations, although Apple has not formally announced it and the M8 Ultra has not yet been released. Baltra’s first AI server chip is expected to use TSMC’s 3-nanometer N3E process and a chiplet design. Broadcom would help design the individual chiplets, while Apple would assemble them into a single package, allowing Apple to keep the overall architecture confidential from partners. Apple is also considering NVLink Fusion to connect the M8 chips. Nvidia introduced the technology in May 2025 to let third-party custom chips communicate at high speed with one another and work alongside Nvidia GPUs. Nvidia says fifth-generation NVLink can fully connect 72 GPUs, with 1,800 GB/s communication speed and up to 130 TB/s of aggregate bandwidth. MediaTek and Marvell were identified as initial partners, while Fujitsu and Qualcomm were also planning integrations. The report recalls Apple’s discontinued Xserve line, which ran from 2002 to 2011 and failed to gain traction in enterprise computing. Apple’s renewed interest may be supported by strong Mac demand: its latest quarterly Mac revenue reached $10.3 billion, up nearly 29% year over year, with market observers attributing much of the growth to bulk purchases by AI companies.