AI Engineer · AI
Are LLM Performance Benchmarks Reliable? — Ashok Chandrasekar & Jason Kramberger, Google
Ask a benchmark harness for 200 queries a second and it may quietly deliver 38, then print results as though it ran 200. Ashok Chandrasekar opens with that experiment, which is why he and Jason Kramberger, both at Google, kept failing to reproduce published numbers. Python's global interpreter lock makes a single process harness CPU bound; the ones they tested capped near 170 and never said so. A thrashing client also inflates the latency it meas

Introductie van de bron.
AI Engineer