• 2 min read
AMD and Cerebras target Nvidia with faster inference
AMD and Cerebras are combining Instinct GPUs with SRAM accelerators for faster, lower-latency AI inference, with a Cerebras Cloud launch planned later this year.

Image: The Register
AMD is pairing its Instinct GPUs with Cerebras Systems' SRAM-based accelerators in a disaggregated platform designed to reduce latency for agentic AI workloads. The companies announced the collaboration on stage during AMD CEO Lisa Su’s Advancing AI keynote Thursday.
How AMD and Cerebras will split inference workloads
The proposed system assigns compute-intensive prompt processing to AMD’s Instinct GPUs, while Cerebras' wafer-scale engines (WSEs) handle memory-intensive token generation. Cerebras uses on-chip SRAM rather than HBM4, giving its accelerators substantially higher memory speed and helping the company reach output rates that often exceed 2,000 tokens per second.
The companies did not provide specific performance figures, but said the combination could increase tokens generated per watt by as much as 5x. The goal is to improve interactivity without sacrificing throughput or increasing costs.
“What you have with Instinct and the Helios rack is you have the leader in performance and memory capacity. And you marry that with our Wafer Scale Engine, which is the leader in SRAM and in memory bandwidth, and that combination allows us to deliver a solution that is unmatched.”
The partnership fills a gap in AMD’s portfolio that helped prompt Nvidia’s $20 billion acquihire of Groq in December. Cerebras' accelerators serve a role similar to Groq 3 LPUs, or Language Processing Units, which Nvidia announced alongside its Vera Rubin rack systems at GTC in March.

Recommended reading
AMD challenges Nvidia with Helios AI racks
For a trillion-parameter model such as Kimi K2.5, Nvidia reportedly needs 2,000 Groq LPUs worth of SRAM. AMD and Cerebras say their approach would require at most a few dozen accelerators.
Cerebras Cloud availability
The combined offering is scheduled to arrive in Cerebras Cloud later this year. AMD’s partnership with Cerebras may also expand beyond this initial platform.
“There are lots of ways to get workload-specific acceleration done, and I think Cerebras has a very interesting technology. It works very well with Helios. The idea of our open ecosystem is frankly that we will work with a number of different companies that may have technology that could be useful.”
Su added: “You can expect that we’re going to do more workload disaggregation going forward.”
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.
via The Register


