Radar van Elk Solutions

AI Engineer · AI

Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab

Your agent scores great on benchmarks. Then real customers use it and it breaks. Felipe Blanes from the Amazon AGI Lab shares what he learned working directly with customers of Nova Act, Amazon's service for building browser agents, from research preview to general availability on AWS. He calls the core problem the benchmark illusion: static evals look great until customers do things nobody expected, and then all you have left is hope. His fix is

Introductie van de bron.

AI Engineer