Search
IDCNOVA

AMD and Cerebras Team Up to Build Ultra-Low-Latency AI Inference System

By: IDCNOVARegion: North America
AMD has entered into a technical partnership with rival chip company Cerebras to develop a disaggregated AI inference solution. The collaboration combines AMD’s new Helios rackscale solution with Cerebras’ Wafer-Scale Engine (WSE), the company’s large-scale AI chip. The integrated system is designed to operate as a single inference workflow, with AMD Helios handling high-performance, scalable throughput for processing prompts and large context windows, while Cerebras delivers ultra-fast, memory-bandwidth-intensive token generation.

The partnership is a direct response to the growing demand for ultra-low latency token generation in AI applications. This trend has already driven Nvidia’s semi-acquisition of Groq, which led to the deployment of the LPX rack featuring Groq LPUs, and Intel’s similar partnership with SambaNova. Cerebras plans to deploy AMD Helios systems in its own data centers, with the joint solution initially available through Cerebras Cloud in the second half of 2026. A broader rollout is expected to follow the cloud launch.

“At Cerebras we build the world’s largest and fastest chip,” said CEO Andrew Feldman, highlighting the company’s large customer base. “They deploy us because we’re blisteringly fast.” He added that AI has evolved from a novelty to a necessity in some domains, and that speed has become critical. “We saw a partnership where we could extend our footprint in ultra-low latency... It’s really something amazing.”

The combined offering positions AMD and Cerebras to compete directly with Nvidia’s NVL72 and Groq LPX systems. By disaggregating the inference pipeline, the solution aims to optimize both throughput and latency, addressing a key bottleneck in real-time AI workloads. This reflects a broader industry shift toward specialized, modular architectures for AI inference, where different hardware components are optimized for distinct tasks within a single workflow.