AMD and Cerebras Team on AI Inference

AMD and Cerebras unveil a new AI inference solution combining AMD Helios GPUs with Cerebras' Wafer-Scale Engine for ultra-low latency and high throughput.

AMD and Cerebras logos side-by-side with abstract AI graphics
AMD and Cerebras join forces for advanced AI inference solutions.
Visual TL;DR
Growing AI Inference DemandsDriver
increasing need for specialized infrastructure in AI inference workloads
From the article 3 mentionsAMD and Cerebras are joining forces to tackle the growing demands of AI inference, announcing a new solution designed for ultra-low latency and high throughput.
AMD + Cerebras TeamCore
joining forces to tackle growing demands of AI inference
From the article 5 mentionsThis collaboration merges AMD's Helios rack-scale systems with Cerebras' Wafer-Scale Engine.
New AI Inference SolutionContext
combining AMD Helios GPUs with Cerebras Wafer-Scale Engine
From the article 7 mentionsAMD and Cerebras are joining forces to tackle the growing demands of AI inference, announcing a new solution designed for ultra-low latency and high throughput.
Disaggregated WorkflowContext
Helios handles high-throughput, Cerebras accelerates token generation
From the article 3 mentionsThe partnership creates a disaggregated inference workflow.
5x Tokens/Sec/WattOutcome
From the articleThe companies claim this combination can deliver up to five times higher tokens per second per watt.
Ultra-Low LatencyEffect
Wafer-Scale Engine accelerates memory-bandwidth-intensive token generation with minimal delay
From the articleAMD and Cerebras are joining forces to tackle the growing demands of AI inference, announcing a new solution designed for ultra-low latency and high throughput.
High ThroughputEffect
AMD Helios systems process prompts and large context windows efficiently
From the articleAMD and Cerebras are joining forces to tackle the growing demands of AI inference, announcing a new solution designed for ultra-low latency and high throughput.
Optimized AI InferenceOutcome
From the article 6 mentionsThis move addresses the increasing need for specialized infrastructure in AI inference, where different workloads have vastly different requirements for speed, capacity, and cost.

AMD and Cerebras are joining forces to tackle the growing demands of AI inference, announcing a new solution designed for ultra-low latency and high throughput. This collaboration merges AMD's Helios rack-scale systems with Cerebras' Wafer-Scale Engine.

StartupHub data

Companies working on this

Profiles of the companies named in this story, with funding and a one-liner from our database.

AMD
Designs and manufactures high-performance microprocessors and graphics processors for computing and data center markets.

The partnership creates a disaggregated inference workflow. AMD Helios will handle high-throughput tasks, processing prompts and large context windows, while the Cerebras Wafer-Scale Engine will accelerate memory-bandwidth-intensive token generation with minimal delay. The companies claim this combination can deliver up to five times higher tokens per second per watt.

This move addresses the increasing need for specialized infrastructure in AI inference, where different workloads have vastly different requirements for speed, capacity, and cost. High-volume tasks might prioritize raw token generation, whereas real-time applications like copilots and agentic workflows demand near-instantaneous responses.

The collaboration aims to optimize these distinct stages independently within a unified workflow. This approach allows for tailored performance without compromising overall efficiency or scalability, creating a differentiated platform for latency-sensitive AI applications.

A New Frontier for Real-Time AI

"AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach," stated Dr. Lisa Su, chair and CEO of AMD. "Together with Cerebras, we are extending that leadership into the most latency-sensitive applications and creating a powerful new platform for real-time agentic AI." This sentiment echoes the strategic importance AMD places on systems like AMD Helios, as discussed in light of recent financial performance.

Andrew Feldman, CEO and co-founder of Cerebras, emphasized the rapid growth in demand for fast inference. "Partnering with AMD gives us an incredible opportunity to bring that performance to even more customers," he said. Cerebras, known for its ambitious chip designs, sees this partnership as a way to broaden the reach of its ultra-fast inference capabilities, building on the momentum discussed around Cerebras Wafer-Scale Engine technology.

The joint solution is expected to be available first through Cerebras Cloud in the second half of 2026. Cerebras plans to deploy AMD Helios systems within its own data centers as part of this initiative.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer