AI Engineer · AI
What's New in Inference Engineering — Philip Kiely, Baseten
TurboQuant reached twenty million people in March, and the memory stock index dipped because everyone assumed the KV cache had just halved. Philip Kiely had published Inference Engineering weeks earlier and watched a technique he had not covered go viral. So his team did the math. Four bit cache instead of eight does double effective bandwidth, but the extra decode computation cuts tokens per second by more than half: unacceptable in a data cente

Introductie van de bron.
AI Engineer