Akshath Tiwari

This is a deep dive on GEPA, short for Genetic-Pareto, a reflective prompt optimizer from Agrawal and colleagues 2025, arXiv 2507.19457, accepted as an ICLR 2026 oral. Its central argument is that a scalar reward is a thin learning signal because it compresses an entire trajectory into a single number, whereas the text of what went wrong carries far more information; so instead of a reinforcement-learning method like GRPO that samples a trajectory, scores it with a reward, and nudges the weights up the policy gradient, GEPA hands the whole execution trace, the reasoning, tool calls, tool outputs, and the evaluation metric’s textual feedback, to a reflector language model that diagnoses the failure in plain language and rewrites the offending prompt, a directed mutation the paper calls Actionable Side Information, the text-optimization analogue of a gradient. The load-bearing second idea is candidate selection. Always mutating the current best candidate collapses into a local optimum after roughly one improvement, so GEPA instead builds an instance-wise Pareto frontier: for each validation instance it finds the best score any candidate achieves and keeps every candidate that ties that best, then prunes strictly dominated candidates and samples the next parent stochastically with probability proportional to how many instances each candidate wins, which means a prompt with a mediocre average survives as long as it is the sole master of one hard instance, preserving every insight discovered in any mutation while still favouring broad winners, the exploration exploitation balance the greedy rule destroys. A system-aware merge step provides crossover by splicing two Pareto-optimal candidates that improved different modules of a compound system into one child that inherits both strengths. The reported results: across its benchmark suite GEPA outperforms GRPO by 6 percent on average and by up to 20 percent while using up to 35 times fewer rollouts, turning a budget of roughly 24000 rollouts into a few hundred, and it beats the leading prompt optimizer MIPROv2 by more than 10 percent, for example plus 12 percent on AIME-2025, with prompts about 33 percent shorter. Per task GEPA scores on Qwen3-8B and GPT-4.1-mini were HotpotQA 62.33 and 69.00, IFBench 38.61 and 52.72, HoVer 52.33 and 51.67, PUPA 91.85 and 94.47, for aggregate gains over the baseline of plus 12.44 percent and plus 14.29 percent. As a preliminary inference-time code search GEPA lifted GPT-4o on NPUEval from 4.25 percent to 30.52 percent mean vector utilization. The honest caveats are that rich textual feedback risks prompt overfitting and bloat so it needs a held-out validation set and length regularization, the reflector must be a strong frontier model because it reasons about reasoning, the failure mode must be expressible in language, and in data-abundant regimes full fine-tuning may still win. Two interactive widgets let the reader flip between greedy and Pareto selection on a shared score matrix to watch greedy discard the one prompt that uniquely cracks a hard instance, and slide a rollout budget to compare reflective versus policy-gradient sample efficiency.