Radar van Elk Solutions

AI Engineer · AI

The End of TCP for AI Clusters — John Ousterhout, Stanford

Split a job across several nodes, let their GPUs compute, then have them exchange a little metadata before the next round. While that exchange happens every GPU sits idle, and the whole cluster waits on the slowest one. Back when a compute phase ran five seconds and the exchange took a few milliseconds, nobody cared. Agentic inference has pushed those compute phases down to milliseconds, so the synchronization now costs about what the work costs.

Introductie van de bron.

AI Engineer