Radar van Elk Solutions

AI Engineer · AI

Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase

The official judge said the agent succeeded 74% of the time. A better verifier said 38%. Miguel González Fernández, tech lead for Browserbase's agent platform, and Corby Rosset, researcher at Microsoft Research, present their research on verifiers for computer-use and web agents. Deterministic evals stopped scaling as agents improved, and the LLM judges bundled with popular web benchmarks turned out to be confidently wrong. Train against them and

Introductie van de bron.

AI Engineer