Cerebras showed GPT-OSS running at over 4,400 tokens per second on CS4 at Hot Chips, and promised 10,000 tokens per second next year with CS5. The interview with CTO Sean Lie on Latent Space at Crisis HQ makes clear why that number matters now.
What used to feel fast at 100 to 200 tokens per second is becoming batch mode. Cerebras is betting that latency, not just throughput, decides who can ship agentic loops and interactive use cases.
Why this matters for AI and startups
CS4 is built on the new Nexus platform. Lie said it delivers twice the power to the wafer, twice the interconnect bandwidth, and half the latency versus the prior generation.
The result is another 2x step. Lie described Cerebras as already the undisputed leader in ultra-fast inference, with CS4 pushing that another two times.
CS5 is built for the same platform and due next year. Lie said it adds another 2x, enabling medium models like GPT-OSS or Gemma at up to 10,000 tokens per second and frontier models like Kimi, DeepSeek and GPT-5 class at up to 5,000 tokens per second.
