AI Engineer · AI
Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave
On an agentic request, somewhere between 80 and 90 percent of the input is identical to the request before it, and prefill is the most expensive thing an inference stack does. That fact, Sitanshu Gupta argues, is why a cached input token is priced so far below a fresh one, and it shapes nearly every choice in the platform he now runs at CoreWeave, four months in. It has to serve two consumption models without forking. Serverless, where you pay pe

Introductie van de bron.
AI Engineer