Radar van Elk Solutions

AI Engineer · AI

Two Bugs That Hid in Plain Sight: A vLLM Debugging Detective Story — Asaf Gardin & Yuval Belfer

Once in roughly a thousand prompts, the model returned gibberish. No crash, no warning, and high confidence, which made it an engineering problem, not a quality one. It happened only in vLLM, only under load, and only with Jamba, AI21's hybrid of attention and Mamba layers. Asaf Gardin and Yuval Belfer could not reproduce it with prompts, so they starved it: dropping vLLM's GPU memory utilization from ninety percent to twenty, at temperature zero

Introductie van de bron.

AI Engineer