Cerebras hits 10,000 tokens per second with CS5

Cerebras showed 4,400 TPS on CS4 and promised 10,000 TPS on CS5, with OpenAI already running its largest model 14x faster.

Cerebras wafer-scale chip powering ultra-fast AI inference up to 10,000 tokens per second
Cerebras CTO Sean Lie at Hot Chips on CS4, CS5 and the shift to ultra-fast inference· Latent Space
Contents(3)

Cerebras showed GPT-OSS running at over 4,400 tokens per second on CS4 at Hot Chips, and promised 10,000 tokens per second next year with CS5. The interview with CTO Sean Lie on Latent Space at Crisis HQ makes clear why that number matters now.

Companies working on this

StartupHub profiles of the companies this article names, with funding and a one-liner from our database.

OpenAI
An artificial intelligence research organization developing and promoting friendly AI for the benefit of humanity.
Groq
$1.0B
Groq develops a high-performance AI inference chip and compiler for ultra-low latency AI applications.
AMD
Designs and manufactures high-performance microprocessors and graphics processors for computing and data center markets.
OpenAI
$13.0B
Artificial intelligence research and deployment company focused on developing advanced AI models like GPT-5.6 and GPT-Live, offering products such as ChatGPT and an API platform.
Cerebras hits 10,000 tokens per second with CS5 - Latent Space
Cerebras hits 10,000 tokens per second with CS5, from Latent Space

What used to feel fast at 100 to 200 tokens per second is becoming batch mode. Cerebras is betting that latency, not just throughput, decides who can ship agentic loops and interactive use cases.

Why this matters for AI and startups

CS4 is built on the new Nexus platform. Lie said it delivers twice the power to the wafer, twice the interconnect bandwidth, and half the latency versus the prior generation.

The result is another 2x step. Lie described Cerebras as already the undisputed leader in ultra-fast inference, with CS4 pushing that another two times.

CS5 is built for the same platform and due next year. Lie said it adds another 2x, enabling medium models like GPT-OSS or Gemma at up to 10,000 tokens per second and frontier models like Kimi, DeepSeek and GPT-5 class at up to 5,000 tokens per second.

For builders, that changes product design. Lie argued real-time agentic loops and reasoning get more capable when you can run thousands of tokens per second, turning offline batch jobs into interactive experiences.

Capacity tells the other story. Lie said Cerebras is sold out of everything it is building, with megawatts deployed strategically and a large share going to OpenAI.

OpenAI is using that capacity internally first. Lie cited incident response and research where extra reasoning and speed have high value, with enterprise access now opening as part of the ultra-fast launch that runs OpenAI's largest model 14 times faster than GPU baselines.

Lie positioned this as co-design, not just acceleration. Models today are designed for a specific Nvidia GPU like B200 or GB200, so even a 14x speedup leaves headroom if model architectures are tuned for wafer-scale.

The competitive read is sharp. Lie called OpenAI's Jalapeno the most exciting Hot Chips reveal, not for its throughput versus Nvidia, but for its AI-first design methodology that built a significantly better GPU faster.

He welcomed Jalapeno plus CS5 as a full fast inference portfolio next year. The combination enables prefill and decode disaggregation and other splits where throughput hardware and latency hardware each do what they do best.

On Groq, Lie was pointed. He said Groq's new LPU numbers were shown only on a non-disaggregated 30 billion parameter model, with no mention of the attention and FFN disaggregation Jensen Huang detailed at GTC for Rubin.

His argument was architectural. A wafer holds about 100 times more SRAM than a single LPU, so a multi-trillion parameter frontier model would need thousands of LPUs just to hold weights. That pushes such designs toward smaller models.

For traditional roadmaps, Lie put Nvidia, AMD, Trainium and TPU on the same treadmill. They are building a better Rubin to deliver more tokens cheaper, which is vital for prompt processing and parallel workloads but not enough for the ultra-fast tier.

What the 10,000 tokens per second push still leaves unanswered

Lie did not share power per token, cost, or yield data for CS4 or CS5. Without those, builders cannot model unit economics for the 10,000 tokens per second tier versus 200 tokens per second batch.

He also acknowledged the packaging problem that will decide this tier. Drawing 3D DRAM or distributed memory on a slide is easy, but power delivery and cooling at wafer scale is hard, which is why Cerebras started its own DRAM stacking program two years ago and leaned on its wafer-scale yield and 3D packaging experience for the backpack power design.

The startup angle others will miss is how heterogeneous this market has become. Inference is no longer one workload. It is prefill, decode, KV cache load, attention and expert balancing, and at multi-gigawatt scale it pays to carve the data center like a chip with different silicon for each slice. That favors vendors who can sell a system, not just a chip.

Lie closed on supply chain risk with a blunt note: most high quality open models are now Chinese, and the hardware to run them is being built in parallel. He said closing that gap will need more than one company or fab.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer

Startups in this story

Profiles for the companies named above.