Date: September 21, 2026
Author: Zhongzhu Zhou
Paper reviewed: GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
Paper authors: Lakshya A Agrawal, Shangyin Tan, Dilara Soylu, Noah Ziems, Rishi Khare, Krista Opsahl-Ong, Arnav Singhvi, Herumb Shandilya, Michael J Ryan, Meng Jiang, Christopher Potts, Koushik Sen, Alexandros G. Dimakis, Ion Stoica, Dan Klein, Matei Zaharia, and Omar Khattab
arXiv: 2507.19457v2, February 14, 2026; first submission July 25, 2025
Reading scope: The complete 96-page version, including algorithms, experimental protocols, evolved prompts and reflection-call counts. Reported measurements below belong to the paper; worked examples and proposed evaluations are identified as this review’s analysis.
1. The contribution in one concrete example
An agent can fail a question even when its underlying language model already knows how to solve each individual step. A first retrieval might find the right person, but its summary drops a relationship. The next query repeats the original question. The final answer generator then has no evidence for the missing relation. Giving the entire trajectory a reward of zero identifies the failure but discards much of its structure.
GEPA turns the trajectory and its diagnostic feedback into material for changing the instructions of one module. A reflection model can propose a rule such as preserving bridge entities in summaries or directing the next search toward missing evidence. The revised instruction is tested, retained when it passes a small comparison, and entered into a population of candidate programs. The optimization therefore happens in readable prompts while the underlying model weights remain fixed.
There are three distinct contributions worth separating. First, the proposal mechanism uses rich feedback instead of only scalar scores. Second, parent selection preserves candidates that are particularly good on some validation examples, avoiding immediate collapse onto one average winner. Third, optional merging combines prompt changes from related branches of a modular system. These mechanisms answer different questions: what to change, where to search next, and how to reuse earlier improvements.
The empirical result is substantial but bounded. On the six tasks in Table 1, GEPA with Qwen3-8B has a macro-average score of 54.85, versus 48.91 for the paper’s GRPO setup. That is 5.94 percentage points, with an AIME-2025 loss alongside five task wins. The famous 35-fold rollout comparison uses the point at which GEPA discovers its best IFBench prompt, not its entire optimization budget and not an all-inclusive dollar cost. Those distinctions are central to evaluating the contribution.

The useful interpretation is that language can carry a reusable task-specific update. It does not establish that prompts universally replace weight training. A model must still understand the feedback, express the required behavior, and carry it out reliably when the prompt is used on unseen inputs.
2. Prerequisites: programs, trajectories and feedback
2.1 A compound system is the optimization unit
A compound language-model system connects multiple modules through a control flow. A retrieval system might have a query generator, a summarizer, another query generator and an answer module. Each module has instructions, a model, an input schema and an output schema. Changing the summarizer can affect everything downstream even when no other prompt changes.
Let denote all module prompts, the collection of model weights, and the final system output on input . Let contain the task’s evaluation metadata, such as a reference answer or required evidence documents. The optimization objective can be written as
The paper’s prompt optimizer searches over with fixed, under a finite rollout budget. A rollout is a system execution, potentially involving several model calls and tools. It is not automatically one model call, one generated sequence of equal length, or one unit of computational cost.
The expectation hides two sources of variation. Inputs vary across tasks, and a stochastic model can produce different trajectories for the same input and prompt. A prompt improvement measured on one trajectory can therefore be accidental. This matters both for the small acceptance minibatch and for the stored validation scores used in selection.
2.2 Execution traces and evaluation traces contain different evidence
An execution trace records what the system did: module inputs, generated text, tool requests and observations. An evaluation trace explains how the evaluator reached its judgment: a missing constraint, an absent supporting document, a compiler diagnostic, or a privacy/quality breakdown. Neither is interchangeable with the final scalar reward.
Consider two failed retrieval trajectories. One asks an irrelevant query; the other retrieves the right evidence but loses a key date during summarization. Their final score may be identical. Their appropriate prompt changes should differ. Reflection can read these distinctions and propose a targeted instruction. The feedback function thus acts as an interface between the task’s measurement machinery and the optimizer.

The diagnosis remains a hypothesis. Seeing that an answer is wrong and a summary is incomplete does not prove that editing the summary is the best intervention. GEPA tests the resulting program to decide whether the proposal is useful. Reflection supplies a structured prior over edits; empirical evaluation still supplies the acceptance signal.
2.3 The difference from a weight update
A simplified policy-gradient identity helps clarify the contrast. If a trajectory is sampled from a differentiable policy , then, subject to the usual regularity conditions,
This identity is background intuition, not the full GRPO objective used in the experiment. GRPO adds group-relative estimation and its training machinery. GEPA instead samples a textual proposal from an implicit distribution
where is textual feedback. There is no computed derivative with respect to prompt tokens and no theorem that a plausible explanation points uphill. The reflection model contributes knowledge acquired during its own pretraining. GEPA uses that knowledge to compress observed failures into instructions and then measures the result.
3. The mutation loop, step by step
3.1 Separate feedback, selection and final evaluation
The generic algorithm divides available development data into a feedback set and a Pareto selection set. In the experiments, predefined training and validation splits serve those roles. The reflection process uses training examples and their diagnostic traces. Validation scores guide parent selection and the choice of the final program. The held-out test set is used to report generalization.
These are three different information channels. Preventing the reflection model from reading validation text does not make validation statistically independent of optimization: the score vector still influences which ancestors produce descendants. It is a development set reused adaptively. The final test must stay separate if its score is to represent performance beyond that search process.
3.2 Algorithm 1: core reflective evolution
The following numbered pseudocode restates the main paper algorithm in this review’s notation. The optional merge proposal is omitted here to keep mutation’s admission rule clear.
- Initialize. Put the original prompt program in archive ; record its parent as empty.
- Score. Evaluate on every selection instance and store its score vector .
- Select. While the rollout budget permits another iteration, sample a parent using Algorithm 2 below.
- Choose a module. Select one module for editing; the reported default cycles through modules in round-robin order.
- Collect feedback. Sample training examples. Execute the parent and retain scores, relevant traces and textual diagnostics.
- Propose. Give the current instruction and feedback to the reflection model. Replace only the chosen module’s prompt in a copy of the parent.
- Compare. Execute the child on the same three examples; compare its average score with the parent’s.
- Admit conditionally. If the child strictly improves the minibatch average, evaluate it on the complete selection set and add the candidate, score vector and lineage to the archive.
- Continue. A rejected edit does not replace its parent. Other retained candidates remain available according to the selection policy.
- Return. When optimization ends, return one archived candidate with the highest mean selection-set score.
The final step is easy to misread. GEPA’s candidate diversity is a search mechanism. The core adaptation procedure does not serve an ensemble that chooses a different prompt for each test example. Per-instance winners help decide which programs to explore, but deployment returns a single selected program.
3.3 Why use a three-example gate?
Full validation is expensive. A cheap local comparison rejects many unpromising edits before they consume an entire validation pass. Using the same examples before and after the edit also removes some between-example variability: a difficult child batch cannot be blamed merely on different sampled questions.
The tradeoff is statistical weakness. For binary outcomes, three examples allow only a few possible score differences. One changed answer can decide admission. For continuous rewards, a small change on a single example can also pass the strict inequality. The gate is a budget heuristic, not a significance test.
An illustrative accounting identity makes the design choice explicit. Suppose there are proposed edits, accepted edits, selection examples, and each mutation requires two fresh batches of size . Ignoring merges and implementation-specific reuse,
The first term initializes the archive, the second tests parent/child pairs, and the third evaluates admitted children. Reflection calls are additional model work. If parent results are cached or a final iteration stops early, the accounting changes. This is a planning model, not a reconstruction of the paper’s precise counters. It explains why admission frequency and selection-set size can matter more than the number of mutations alone.
3.4 Why change only one module?
A single-module edit narrows the proposal space and makes a successful change easier to interpret. It also leaves the surrounding interfaces intact. Editing every module at once could fix a coordinated failure faster, but it creates many explanations for a score change and can destroy previously useful behavior.
Round-robin scheduling avoids needing a learned fault-localization policy. Its cost is spending attention on modules that may not be the current bottleneck. A useful future comparison would hold reflection and evaluation budgets fixed while selecting modules by failure frequency or marginal intervention value. That would test whether smarter module choice improves the optimizer, without conflating it with a larger search budget.
4. What instance-wise selection actually does
4.1 Deriving the sampling distribution
Let be candidate ‘s score on validation example . For each example, first compute the maximum observed score and the set of candidates attaining it:
Take the union of these winner sets. Remove candidates dominated by another candidate in the union. Here domination means no lower score on any example and a strictly higher score on at least one. Denote the surviving per-example sets by . Then count memberships and normalize:
The denominator counts winning memberships, including ties, rather than merely the number of validation examples. A candidate can receive several tickets in the sampling pool by excelling on different examples. A specialist that loses in average score can therefore continue generating descendants.
4.2 Algorithm 2: parent selection
- For every selection example, find the highest candidate score.
- Collect all candidates tied at that score into the example’s winner set.
- Form the union of all winner sets.
- Remove dominated candidates from this union and from the corresponding winner sets.
- Count how many surviving winner sets contain each candidate.
- Sample a parent with probability proportional to that count.

In Figure 3, A wins the first example, B wins the first and second, and C wins the third and fourth. D wins none and is also dominated. Their counts are , giving probabilities . If instead we sampled an example uniformly and then sampled uniformly among its tied winners, the probabilities would be . The two procedures differ. Keeping that detail straight matters when interpreting the algorithm’s exploration bias.
4.3 This is a restricted set, not the entire Pareto frontier
The published Algorithm 2 first takes coordinate-wise winners and then removes dominated candidates. A general nondominated candidate need not win any coordinate. Consider
None dominates another. C has the best average score, yet A and B supply all coordinate maxima, so C receives no parent-sampling mass under this rule. C can still be returned by final average-based selection if it is already in the archive. Search eligibility and final selection are separate operations.

This is not necessarily a defect: preserving specialists may help discover complementary rules. It is a real inductive bias. A population of balanced, near-best candidates can contain useful starting points that the rule never samples. An exploration mixture that assigns a small probability to the full archive would be a natural test; its benefit is a hypothesis rather than a result reported by the paper.
The rule also depends on measurement quality. With stochastic outputs, one lucky score can make a candidate an instance winner. With many validation dimensions, dominance may become rare. Repeating selected evaluations or grouping examples into stable task categories could change both cost and diversity. Those interventions should be assessed with a fixed total budget, since otherwise an apparent algorithmic gain may simply reflect more evaluation.
5. Merging useful instructions across branches
5.1 A program-level crossover
Suppose two descendants share an ancestor. One improves query generation; the other improves summarization. Copying both changed modules into a new candidate can combine lessons without asking a reflection model to rediscover them. This is the motivation for system-aware merge in Appendix D.1. It is an optional proposal mechanism and is used sparsely, with at most five merges in the reported settings.
The prose emphasizes complementary changes. The detailed pseudocode is more permissive than requiring completely disjoint sets of changed modules. Its desirability check returns true when it finds at least one module changed in exactly one branch. Conflicts elsewhere are allowed and receive a resolution rule. That distinction follows from the paper’s Algorithms 3–4 themselves.
5.2 Algorithm 3: merge proposal as specified in the appendix
- Sample two distinct candidate indices ; skip a direct ancestor/descendant pairing.
- Find their common ancestors. For a chosen ancestor , skip a triple already attempted.
- Skip the combination if the ancestor’s aggregate score is greater than the smaller parent score. Both parents must therefore be at least as good as that ancestor by this comparison.
- Require at least one module where exactly one parent’s prompt differs from the ancestor’s prompt.
- Start a proposed child from the ancestor.
- For each module, copy a one-sided change from the branch that made it.
- If both branches changed a module differently, choose the version from the parent with the higher aggregate score; break ties randomly. Shared changes/default cases retain the matching parent version.
- Return the proposed program and lineage for subsequent evaluation. A constructed merge is not evidence of a performance gain.

Selecting a module by its parent’s aggregate performance is cheap. It does not estimate that module’s isolated value. A high-performing parent might owe its advantage to another module. In Figure 5, choosing answer prompt over from parent scores alone is a heuristic.
5.3 Why two good edits can make a bad combination
Let be the expected system score with query prompt and summary prompt . Relative to an ancestor , define the two single-edit improvements and their interaction:
Rearranging terms gives
Even if both single-edit gains are positive, a sufficiently negative interaction makes the merge worse. A summarizer that removes repeated context may conflict with a query generator that expects that context to be repeated explicitly. This is a system interface problem, not a contradiction of the individual parent scores.
A conservative alternative is to compare a merged candidate on a fresh diagnostic batch before committing a full selection evaluation. A more expensive alternative measures module substitutions across several contexts. Both spend evaluations to understand interactions. The paper’s simple crossover offers a cheaper starting point, but its empirical regressions make this tradeoff concrete.
6. Reading the benchmark evidence without mixing protocols
6.1 Six tasks do not measure one capability
Table 1 combines multi-hop question answering, evidence retrieval, instruction following, privacy-preserving assistance and mathematics. Their macro-average is useful as a compact summary, but it assigns equal weight to distinct metrics and datasets. A one-point gain on a privacy/utility score is not necessarily as valuable as a one-point gain in exact math accuracy.
The following compact transcription preserves the relevant comparison. All numbers are paper-reported scores on a percentage scale; the final row is the reported six-task aggregate.
| Task | Qwen baseline | GRPO | GEPA | GEPA + merge |
|---|---|---|---|---|
| HotpotQA | 42.33 | 43.33 | 62.33 | 64.33 |
| IFBench | 36.90 | 35.88 | 38.61 | 28.23 |
| HoVer | 35.33 | 38.67 | 52.33 | 51.67 |
| PUPA | 80.82 | 86.66 | 91.85 | 86.26 |
| AIME-2025 | 27.33 | 38.00 | 32.00 | 32.00 |
| LiveBench Math | 48.70 | 51.26 | 51.95 | 51.95 |
| Aggregate | 45.23 | 48.91 | 54.85 | 52.40 |
The most convincing gains are on HotpotQA and HoVer, where prompt-level coordination and missing-evidence feedback directly match the method’s strengths. The AIME exception is also informative. Better instructions do not necessarily supply the mathematical competence that weight adaptation can improve. Neither direction can be generalized from six tasks to all agent workloads.

6.2 What the splits and scores mean
HotpotQA and HoVer use 150 training, 300 validation and 300 test examples each. The HoVer system is evaluated on retrieving the complete supporting-document set. It should not be described as a comprehensive factual-veracity benchmark for every final response. The HotpotQA setup attaches an answer module to a multi-hop retrieval system; module-level feedback can expose which supporting documents remain missing.
IFBench uses 150 training and 300 validation examples from IF-RLVR, with 294 IFBench test examples covering new constraints. The two-stage system generates an answer and then revises it for compliance. This offers a meaningful shift in constraint types, but the prompt examples later in the appendix show why transfer is not automatic: a rule that was helpful in observed cases can become an unconditional instruction that is wrong elsewhere.
For AIME, 90 problems from 2022–2024 are divided into 45 training and 45 validation problems. The 2025 test consists of 30 distinct problems, each sampled five times. There are 150 generations, not 150 independent problems. Repeated generations help estimate stochastic success on the same problems; they do not create a larger independent sample of mathematical topics. LiveBench Math uses 368 collected problems, shuffled with seed zero and divided into thirds.
PUPA has 111 training, 111 validation and 221 test examples. Its architecture uses a trusted rewriter, an external model and a trusted final responder. The optimization balances utility and private-information exposure through the task’s feedback. A high aggregate score does not imply that every sensitive attribute is protected, nor that utility is preserved equally for every query class.
6.3 Baselines and their scope
The Qwen3-8B GRPO comparison uses compound-system LoRA training: rank 16, alpha 64, dropout 0.05, a group of 12 outputs, four instances per step and 500 steps, totaling 24,000 rollouts. Its validation is periodic. These details define a specific practical baseline, rather than the strongest possible RL system under every compute allocation.
The supplementary full-parameter GRPO experiment is confined to a two-hop HoVer setting. It is useful additional evidence, but cannot silently expand the six-task LoRA comparison into a six-task claim about all full-parameter training. Likewise, changing optimizer hyperparameters, model scale or reward design might change the ranking.
MIPROv2 provides a relevant prompt-search baseline, including a no-demonstration variant for GPT-4.1 Mini. Trace/OptoPrime and TextGrad are adapted to the same compound architecture and evaluation interfaces. The paper notes that module-specific feedback is not supported in the same way by those alternatives. Consequently, the end-to-end comparison measures the systems as configured, while a clean causal claim about the selection rule needs the separate ablation.
The common inference settings also matter: Qwen uses temperature 0.6, top-p 0.95 and top-k 20; GPT-4.1 Mini uses the dated 2025-04-14 model at temperature 1.0. A 16,384-token context limit is reported. These are historical experiment conditions. Performance should not be presented as a measurement of whichever API endpoint happens to carry a similar name today.
7. Cost: distinguish discovery, optimization and serving
7.1 Reconstructing the 35-fold statement
On Qwen IFBench, the paper reports that GEPA finds its best prompt after 678 rollouts. GRPO’s training budget is 24,000. The ratio is
But Table 1 gives GEPA’s total IFBench optimization budget as 3,593. Comparing total budgets instead yields
These ratios answer different questions. The first describes when the ultimately selected prompt was discovered in that run. Without a stopping criterion that identifies it online, one cannot assume the rest of the search could have been safely skipped. The second compares the reported full budgets, still using rollout counts rather than time or dollars.

Across the six tasks, the reported mean GEPA budget is 3,936 versus 24,000 for GRPO, a ratio of about 6.10. That average compresses substantial variation across tasks. It also leaves the differing costs of optimizer operations outside a single homogeneous unit.
7.2 Reflection calls are finite, but not free
Appendix N makes reflection overhead partly visible. For Qwen3-8B, the counts are 64 for HotpotQA, 17 for IFBench, 50 for HoVer, 38 for PUPA, 90 for AIME and 38 for LiveBench Math. The corresponding GPT-4.1 Mini counts are 69, 21, 92, 46, 24 and 34. These are much smaller than the reported rollout totals, but the input to a reflection call can contain long traces and prompts.
A useful accounting decomposition is
For RL, gradient computation, optimizer state, training/inference synchronization and hardware utilization also enter the comparison. Counting calls cannot price those terms. For an API-based prompt optimizer, prompt and completion tokens, retries, tool execution and latency all matter. A fair report should expose several cost axes rather than converting all of them into a single rollout multiplier.
The appendix gives historical spending for the GPT-4.1 Mini experiments, including USD 86 for GEPA and USD 67 for GEPA with merge. These are costs of those reported experiments, not present-day price quotations or a controlled statement that merging always saves money. Different trajectories of search can consume different resources even under similar budget policies.
7.3 Optimization cost must be amortized
Prompt search spends resources before deployment. A longer evolved instruction may also add tokens on every future request; a shorter instruction may save them. For two candidate systems A and B, a simple workload model is
Here is the number of served requests, is setup/search cost and is per-request cost. The break-even expression is useful when A has higher setup cost but lower serving cost. If the signs differ, or quality is not comparable, that interpretation fails. Caching and batching can also make the linear model approximate.
As an explicitly hypothetical example, a USD 120 optimization investment that saves USD 0.002 per request breaks even after 60,000 requests. These values are illustrative, not GEPA measurements. The point is to connect search improvements to the workload where they will be used. A prompt improvement for a one-off question and a prompt improvement for a service handling millions of requests are different economic decisions.
8. Transfer, selection ablations and the appendix’s lessons
8.1 Transfer offers evidence of reusable instructions
On GPT-4.1 Mini, Table 2 reports a baseline aggregate of 53.03, GEPA at 65.22 and GEPA with merge at 66.36. Prompts optimized entirely with Qwen and transferred unchanged to GPT-4.1 Mini achieve 62.03, a nine-point improvement over the Mini baseline. This is evidence that useful instructions can survive a model change in these tasks.
It does not show that every evolved rule is model-independent. Qwen and GPT have related language capabilities; the test tasks and program architecture remain the same. Transfer across languages, tool schemas, model generations or feedback distributions would be a stronger and different experiment. Nor does the result mean the Qwen-derived prompt is always as good as directly optimizing for the target model.
8.2 Merge exposes a real failure boundary
The Qwen IFBench score falls from 38.61 without merge to 28.23 with merge, a loss of 10.38 points. On GPT-4.1 Mini, merge improves IFBench from 52.72 to 55.95 and HoVer from 51.67 to 56.67, while HotpotQA falls from 69.00 to 65.67. A method can have a positive aggregate effect on one model and a negative one on another.

This also reveals a small but important reporting issue: Table 1’s broad caption suggests both GEPA variants beat GRPO except on AIME, but its Qwen IFBench merge cell is below GRPO. The detailed cells should govern the interpretation. The result supports testing merge as an optional setting rather than assuming crossover is always an upgrade.
8.3 The selection ablation uses four tasks
Table 3 replaces GEPA’s selection policy with greedy selection or a beam of four while keeping the evolutionary harness. The four non-math tasks have averages 48.84 for the initial system, 54.89 for greedy selection, 53.95 for beam selection and 61.28 for GEPA. This is direct evidence that retaining instance specialists matters in the tested setup.

The ablation makes a stronger causal argument about selection than comparing completely different optimization packages. It does not isolate every interaction: the usefulness of candidate diversity depends on the mutation mechanism, reward structure and search budget. A different reflection model or noisier metric could change the preferred parent policy.
8.4 Prompt examples show both abstraction and overgeneralization
The HotpotQA prompts encourage retaining bridge entities and targeting evidence missing after the first hop. These are understandable procedural lessons. Other evolved prompts contain specific names, facts and illustrative cases from observed feedback. The learned object is therefore a mixture of reusable policy and task-specific content, not a pure abstract algorithm.
The Qwen IFBench examples in Appendix L include an unconditional instruction to repeat the user’s query. That may serve some observed repetition constraints but can conflict with unrelated output restrictions. This is a concrete example of a local observation becoming an overly broad rule. It does not by itself prove the cause of the merge regression; a causal explanation would require controlled prompt-component experiments that the displayed examples do not provide.
PUPA prompts progressively remove or generalize more identifying context. This is consistent with stronger privacy protection under the evaluation metric. It also raises a utility question: removing a location, date or numerical parameter can make some legitimate requests impossible to answer precisely. The trust boundary matters too. Sensitive reasoning retained by a trusted module is different from information sent to the external model, so a review should track what crosses that boundary rather than judge privacy from a prompt’s wording alone.
8.5 Inference-time kernel search is a separate setting
For kernel generation, the paper also uses GEPA to search directly on the target task collection. In that setting, training/selection/target tasks can coincide by design. This is useful inference-time optimization, but it is not the same held-out adaptation claim as Tables 1–2.
On NPUEval, the reported GPT-4o baseline score is 4.25, a selected single prompt reaches 26.85, and the Pareto result reported in that experiment is 30.52. A hardware-utilization score is not automatically a speedup multiplier. The CUDA evaluation uses 35 representative KernelBench tasks on a V100, and its fast- metric counts correct kernels exceeding a specified speedup threshold. Those conditions limit extrapolation to other devices and workloads.
The adversarial-prompt example on AIME is another separate measurement. The paper reports a large decline in judged success, but notes that many outputs use a literal answer placeholder. Formatting and parsing failures therefore contribute to the result. It is safer to interpret it as a failure of the evaluated answering pipeline than as clean evidence that the underlying mathematical reasoning ability has disappeared.
9. A stronger evaluation design for practical adoption
The paper motivates a useful development procedure, but a production decision needs more than a higher validation maximum. This section proposes an evaluation design; it does not describe an experiment performed for this review.
First, preserve a locked final test set grouped by the natural sampling unit: question, user request or task family. If a problem is sampled multiple times, keep those repetitions together. Second, track both aggregate utility and disaggregated failures. A privacy task needs separate leakage and task-completion measurements; a constrained-generation task needs semantic correctness and each constraint family; a kernel task needs correctness and latency separately.
Third, compare accepted candidates using paired evaluations. For a fixed parent and child evaluated on independent held-out examples, define
The standard error estimated from paired differences is
Pairing can reduce variance when parent and child succeed on similar examples. It does not remove stochastic generation noise or repair adaptive reuse of the same holdout. A bootstrap should resample independent problem groups rather than individual repeated generations. The three-example mutation gate is too small to substitute for this final assessment.
There is a useful theoretical warning here. For candidates fixed independently of a holdout of size , bounded scores permit a union-bound statement of the form
The steps are a concentration inequality for each fixed candidate followed by a union bound over . Solving for a sufficient at failure probability gives
This is not a guarantee for GEPA’s adaptively generated archive on its reused selection set. Later candidates depend on earlier validation results. Applying the fixed-candidate formula there without accounting for that dependence would be misleading. It is a reason to use a fresh locked evaluation after search, not a confidence certificate for the internal maximum.
Algorithm 4: a proposed decision procedure
- Freeze the task definition, split rules, model version and resource budget before search.
- Run prompt optimization using only development feedback and selection scores.
- Log rollout counts, reflection tokens, total tokens, elapsed time and cost independently.
- Freeze the selected prompt program before touching the final test set.
- Compare it with the baseline and a no-merge variant using paired problem-level evaluation.
- Report confidence intervals and failure subgroups alongside the macro-average.
- Decide whether the quality gain and amortized serving cost justify adoption; preserve the prior prompt for rollback.
This procedure tests the behavior of a selected system. It is deliberately distinct from choosing ever more prompts until a reused test score looks attractive. A rejected result is still useful if it identifies a failed task family or a cost regime where prompt optimization does not pay off.
10. Limitations and failure boundaries
Feedback quality bounds what can be learned. Rich diagnostics can reveal exactly what to fix, but they may require reference documents, trusted labels or expensive judges. A deployment with only delayed user satisfaction has a different information budget. Reflection can also turn an incorrect evaluator explanation into a persistent rule.
Prompt expressiveness is finite. Some shortcomings are coordination problems; others are missing capability, unfamiliar knowledge or unreliable computation. Instructions cannot guarantee that a weak model suddenly performs a long derivation correctly. The AIME result is consistent with this boundary, though it does not isolate its cause.
Search is vulnerable to noise and adaptive overfitting. Small-batch acceptance, per-instance winner selection and repeated validation all favor observed successes. They can amplify luck. The paper’s point estimates are informative, but not a substitute for repeated optimization seeds or uncertainty over the selected system’s performance.
Modularity does not imply independent modules. Input schemas can stay unchanged while the meaning and detail of intermediate text change. A prompt merge can create a semantic interface mismatch. The Qwen IFBench regression is a practical reminder to measure merged candidates independently.
Portability is demonstrated only within a measured scope. The Qwen-to-Mini result is valuable evidence, but does not establish robustness to different languages, tool schemas, model families or long-term API drift. A transferred prompt remains an artifact tied to assumptions about its surrounding system.
Reported efficiency depends on the unit and stopping rule. Rollouts vary in token count, model calls and external work. A retrospective best-prompt point differs from a deployable early-stopping policy. Historical API bills and selected device benchmarks should not be generalized into universal cost multipliers.
11. Independent critical analysis
11.1 The strongest idea is preserving informative feedback
My main takeaway is not the headline contest with RL. It is the interface between execution, evaluation and future instructions. Many agent systems already produce rich evidence of why they failed, then discard it after computing a metric. GEPA offers a concrete way to turn that information into a reusable program change.
This interpretation also clarifies its dependency. Better feedback can improve the optimizer without changing selection or reflection. A future study should vary feedback richness while holding the proposal model, rollout budget and dataset fixed: scalar only, generic explanation, module-specific diagnosis, and reference-rich diagnosis. That would measure how much improvement comes from the optimizer and how much comes from its information channel.
11.2 Selection deserves a more precise name and ablation
The instance-winner rule is more selective than retaining the whole Pareto frontier. Its deliberate preference for specialists is plausible, and the four-task ablation supports it. But the balanced-candidate counterexample suggests an important missing comparison: mix specialist sampling with mean-score sampling or uniform archive exploration.
Such an experiment should measure retained diversity, accepted-edit rate and final held-out utility under the same budget. Otherwise a larger candidate population or additional scoring could masquerade as a better policy. The relevant question is whether a given exploration bias discovers useful general rules efficiently, not whether one frontier visualization looks more diverse.
11.3 Merge’s disappointing cases are scientifically useful
The mismatch between the merge prose and detailed pseudocode, together with model-dependent regressions, argues for reporting the exact variant. The paper permits partially overlapping edits and resolves conflicts using global parent performance. That is a stronger assumption than simply combining independent improvements.
An informative next experiment would separate strictly disjoint merges from conflict-resolving merges, and report their acceptance and test performance separately. A second comparison would replace whole-parent score with a targeted module substitution estimate. These proposals follow from the interaction derivation; they are not findings already established by the paper.
11.4 Prompt optimization and weight training can cooperate
The title correctly says “can.” GEPA’s evidence demonstrates that fixed-weight prompt search can outperform a particular RL setup in several useful settings. It does not rank all methods under all resources. A practical sequence could first remove interface and instruction errors, then train weights on the remaining capability failures, and finally re-optimize prompts for the adapted model.
The potential drawback is a moving objective: every weight update can change which prompt rules are useful. Joint optimization therefore needs a clean evaluation protocol and a cost ledger. GEPA makes prompt programs a tractable object to optimize; it does not remove the need to decide where training, retrieval, architecture changes or human intervention are better investments.
12. Conclusion
GEPA is a persuasive demonstration that an agent’s failed executions contain more learnable structure than their final scalar rewards. Reflection proposes readable instruction changes, instance-wise selection preserves useful specialists, and optional merging reuses improvements across branches. The method’s strongest empirical wins align with tasks where coordination and interpretable feedback matter.
The careful reading is equally valuable: the selected system is one prompt program; instance winners are not the complete Pareto frontier; merge can regress; a best-discovery rollout ratio is not total cost; and repeated samples are not independent test problems. Keeping those distinctions visible makes the result more useful, not less impressive. For an agent workflow with informative diagnostics and repeated use, GEPA provides a concrete optimization strategy worth evaluating under a locked test set and an explicit resource budget.
References and figure provenance
- Agrawal et al. GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning, arXiv v2. Main Algorithms 1–2, Tables 1–3, and Appendices D–N provide the technical and experimental evidence used here.
- Complete GEPA paper PDF. Page 6: main algorithms; page 8: six-task tables; page 24: merge pseudocode; page 96: reflection-call counts. Appendix L contains the evolved prompt examples.
- ICLR 2026 oral program. Venue confirmation.
- Official GEPA project. Resource link for readers.
Figures 1–5 are original explanatory diagrams or synthetic examples by Zhongzhu Zhou. Figures 6–9 are original visualizations of the paper’s tabulated values; their captions identify the relevant table and comparison. The probability example, merge interaction derivation, accounting equations and proposed evaluation procedure are this review’s analysis. No new benchmark or reproduction result is claimed.