AI Engineer · AI
Beating RL With Reflection: GEPA and Optimize Anything — Lakshya A. Agrawal, GEPA
After one round of reflection on just three examples, GEPA doubled the gains that the RL algorithm GRPO reached after 25,000 rollouts. Lakshya A. Agrawal, creator of GEPA and a PhD student at UC Berkeley's Sky Computing Lab, explains why. RL squeezes a whole rollout down to a single score. GEPA instead has a model read the full trace, including chains of thought, tool calls and error messages, and write a better prompt. A Pareto pool of candidate

Introductie van de bron.
AI Engineer