Cerebras showed GPT-OSS running at over 4,400 tokens per second on CS4 at Hot Chips, and promised 10,000 tokens per second next year with CS5. The interview with CTO Sean Lie on Latent Space at Crisis HQ makes clear why that number matters now.
Companies working on this
StartupHub profiles of the companies this article names, with funding and a one-liner from our database.
What used to feel fast at 100 to 200 tokens per second is becoming batch mode. Cerebras is betting that latency, not just throughput, decides who can ship agentic loops and interactive use cases.
