AMD and Cerebras Team on AI Inference

AMD and Cerebras unveil a new AI inference solution combining AMD Helios GPUs with Cerebras' Wafer-Scale Engine for ultra-low latency and high throughput.

7 min read
AMD and Cerebras logos side-by-side with abstract AI graphics
AMD and Cerebras join forces for advanced AI inference solutions.

Visual TL;DR. Growing AI Inference Demands drives AMD + Cerebras Team. AMD + Cerebras Team creates New AI Inference Solution. New AI Inference Solution features Disaggregated Workflow. Disaggregated Workflow enables Ultra-Low Latency. Disaggregated Workflow enables High Throughput. New AI Inference Solution achieves 5x Tokens/Sec/Watt. 5x Tokens/Sec/Watt leads to Optimized AI Inference. Ultra-Low Latency contributes to Optimized AI Inference. High Throughput contributes to Optimized AI Inference.

  1. Growing AI Inference Demands: increasing need for specialized infrastructure in AI inference workloads
  2. AMD + Cerebras Team: joining forces to tackle growing demands of AI inference
  3. New AI Inference Solution: combining AMD Helios GPUs with Cerebras Wafer-Scale Engine
  4. Disaggregated Workflow: Helios handles high-throughput, Cerebras accelerates token generation
  5. Ultra-Low Latency: Wafer-Scale Engine accelerates memory-bandwidth-intensive token generation with minimal delay
  6. High Throughput: AMD Helios systems process prompts and large context windows efficiently
  7. 5x Tokens/Sec/Watt: combination delivers significantly higher tokens per second per watt
  8. Optimized AI Inference: addresses increasing need for specialized infrastructure in AI inference
Visual TL;DR
Visual TL;DR, startuphub.ai Growing AI Inference Demands drives AMD + Cerebras Team. AMD + Cerebras Team creates New AI Inference Solution. New AI Inference Solution achieves 5x Tokens/Sec/Watt. 5x Tokens/Sec/Watt leads to Optimized AI Inference drives creates achieves leads to Growing AI Inference Demands AMD + Cerebras Team New AI Inference Solution 5x Tokens/Sec/Watt Optimized AI Inference From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Growing AI Inference Demands drives AMD + Cerebras Team. AMD + Cerebras Team creates New AI Inference Solution. New AI Inference Solution achieves 5x Tokens/Sec/Watt. 5x Tokens/Sec/Watt leads to Optimized AI Inference drives creates achieves leads to Growing AIInference Demands AMD + CerebrasTeam New AI InferenceSolution 5xTokens/Sec/Watt Optimized AIInference From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Growing AI Inference Demands drives AMD + Cerebras Team. AMD + Cerebras Team creates New AI Inference Solution. New AI Inference Solution achieves 5x Tokens/Sec/Watt. 5x Tokens/Sec/Watt leads to Optimized AI Inference drives creates achieves leads to Growing AI Inference Demands increasing need for specializedinfrastructure in AI inference workloads AMD + Cerebras Team joining forces to tackle growing demandsof AI inference New AI Inference Solution combining AMD Helios GPUs with CerebrasWafer-Scale Engine 5x Tokens/Sec/Watt combination delivers significantly highertokens per second per watt Optimized AI Inference addresses increasing need for specializedinfrastructure in AI inference From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Growing AI Inference Demands drives AMD + Cerebras Team. AMD + Cerebras Team creates New AI Inference Solution. New AI Inference Solution achieves 5x Tokens/Sec/Watt. 5x Tokens/Sec/Watt leads to Optimized AI Inference drives creates achieves leads to Growing AIInference Demands increasing need forspecializedinfrastructure in… AMD + CerebrasTeam joining forces totackle growingdemands of AI… New AI InferenceSolution combining AMDHelios GPUs withCerebras… 5xTokens/Sec/Watt combinationdeliverssignificantly… Optimized AIInference addressesincreasing need forspecialized… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Growing AI Inference Demands drives AMD + Cerebras Team. AMD + Cerebras Team creates New AI Inference Solution. New AI Inference Solution features Disaggregated Workflow. Disaggregated Workflow enables Ultra-Low Latency. Disaggregated Workflow enables High Throughput. New AI Inference Solution achieves 5x Tokens/Sec/Watt. 5x Tokens/Sec/Watt leads to Optimized AI Inference. Ultra-Low Latency contributes to Optimized AI Inference. High Throughput contributes to Optimized AI Inference drives creates features enables enables achieves leads to contributes to contributes to Growing AI Inference Demands increasing need for specializedinfrastructure in AI inference workloads AMD + Cerebras Team joining forces to tackle growing demandsof AI inference New AI Inference Solution combining AMD Helios GPUs with CerebrasWafer-Scale Engine Disaggregated Workflow Helios handles high-throughput, Cerebrasaccelerates token generation Ultra-Low Latency Wafer-Scale Engine acceleratesmemory-bandwidth-intensive tokengeneration with minimal delay High Throughput AMD Helios systems process prompts andlarge context windows efficiently 5x Tokens/Sec/Watt combination delivers significantly highertokens per second per watt Optimized AI Inference addresses increasing need for specializedinfrastructure in AI inference From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Growing AI Inference Demands drives AMD + Cerebras Team. AMD + Cerebras Team creates New AI Inference Solution. New AI Inference Solution features Disaggregated Workflow. Disaggregated Workflow enables Ultra-Low Latency. Disaggregated Workflow enables High Throughput. New AI Inference Solution achieves 5x Tokens/Sec/Watt. 5x Tokens/Sec/Watt leads to Optimized AI Inference. Ultra-Low Latency contributes to Optimized AI Inference. High Throughput contributes to Optimized AI Inference drives creates features enables enables achieves leads to contributes to contributes to Growing AIInference Demands increasing need forspecializedinfrastructure in… AMD + CerebrasTeam joining forces totackle growingdemands of AI… New AI InferenceSolution combining AMDHelios GPUs withCerebras… DisaggregatedWorkflow Helios handleshigh-throughput,Cerebras… Ultra-Low Latency Wafer-Scale Engineacceleratesmemory-bandwidth-int High Throughput AMD Helios systemsprocess prompts andlarge context… 5xTokens/Sec/Watt combinationdeliverssignificantly… Optimized AIInference addressesincreasing need forspecialized… From startuphub.ai · The publishers behind this format

AMD and Cerebras are joining forces to tackle the growing demands of AI inference, announcing a new solution designed for ultra-low latency and high throughput. This collaboration merges AMD's Helios rack-scale systems with Cerebras' Wafer-Scale Engine.

The partnership creates a disaggregated inference workflow. AMD Helios will handle high-throughput tasks, processing prompts and large context windows, while the Cerebras Wafer-Scale Engine will accelerate memory-bandwidth-intensive token generation with minimal delay. The companies claim this combination can deliver up to five times higher tokens per second per watt.

This move addresses the increasing need for specialized infrastructure in AI inference, where different workloads have vastly different requirements for speed, capacity, and cost. High-volume tasks might prioritize raw token generation, whereas real-time applications like copilots and agentic workflows demand near-instantaneous responses.

The collaboration aims to optimize these distinct stages independently within a unified workflow. This approach allows for tailored performance without compromising overall efficiency or scalability, creating a differentiated platform for latency-sensitive AI applications.

A New Frontier for Real-Time AI

"AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach," stated Dr. Lisa Su, chair and CEO of AMD. "Together with Cerebras, we are extending that leadership into the most latency-sensitive applications and creating a powerful new platform for real-time agentic AI." This sentiment echoes the strategic importance AMD places on systems like AMD Helios, as discussed in light of recent financial performance.

Andrew Feldman, CEO and co-founder of Cerebras, emphasized the rapid growth in demand for fast inference. "Partnering with AMD gives us an incredible opportunity to bring that performance to even more customers," he said. Cerebras, known for its ambitious chip designs, sees this partnership as a way to broaden the reach of its ultra-fast inference capabilities, building on the momentum discussed around Cerebras Wafer-Scale Engine technology.

The joint solution is expected to be available first through Cerebras Cloud in the second half of 2026. Cerebras plans to deploy AMD Helios systems within its own data centers as part of this initiative.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.