Latent Space · AI
The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
From pushing inference beyond 4,000 tokens per second to working with OpenAI on a new generation of ultra-fast AI infrastructure, Cerebras is betting that speed doesn’t just make models faster, it makes entirely new kinds of AI possible. In this episode, Cerebras co-founder and CTO Sean Lie joins swyx and Vibhu fresh off Hot Chips to unpack CS4, preview CS5, and explain why 100–200 tokens per second may soon feel like “batch mode.” We go deep on

Introductie van de bron.
Latent Space