AI Engineer · AI
Large clusters for small models — Daniel Svonava, Superlinked
A single mid range GPU can turn half a million tokens per second into embeddings in the low tens of milliseconds, where a managed endpoint costs orders of magnitude more and takes hundreds. Daniel Svonava calls embeddings the no brainer entry point. His real subject is what comes after. A small model fits on one GPU two or three generations old, and for a specific task it is now at or beyond the frontier, which is flattening while small open mode

Introductie van de bron.
AI Engineer