AI Engineer · AI
Evals in AI: A Deep Dive — Tejas Kumar, IBM
A customer support answer passes a test because it contains the word “cannot,” even though it approves a return outside the store's policy. Tejas Kumar builds that failure live to show where ordinary assertions and fuzzy matching stop being useful. The example grows into an LLM judge, then a dataset of customer scenarios, agent responses, and human verdicts. Along the way, the judge starts favoring an answer generated by its own model family. Add

Introductie van de bron.
AI Engineer