Formula

Evo-GDPO Clipped Surrogate Policy Objective

The training objective for Evo-GDPO is formalized as a PPO-style clipped surrogate objective driven by the decoupled, batch-normalized aggregated advantage A^iΣ\widehat{A}_i^\Sigma and regularized by token-level KL divergence:

LIIIEvo-GDPO(θ)=−E[1G∑i=1G1∣hi∣∑j=1∣hi∣min⁡(ρi,j(θ)A^iΣ,clip(ρi,j(θ),1−ϵclip,1+ϵclip)A^iΣ)]+βKLE[KL(pθ∥pref)]\mathcal{L}_{\text{III}}^{\text{Evo-GDPO}}(\theta) = - \mathbb{E} \left[ \frac{1}{G} \sum_{i=1}^G \frac{1}{|\mathbf{h}_i|} \sum_{j=1}^{|\mathbf{h}_i|} \min \left( \rho_{i,j}(\theta) \widehat{A}_i^\Sigma, \text{clip}(\rho_{i,j}(\theta), 1 - \epsilon_{\text{clip}}, 1 + \epsilon_{\text{clip}}) \widehat{A}_i^\Sigma \right) \right] + \beta_{\text{KL}} \mathbb{E} [ \text{KL}(p_\theta \parallel p_{\text{ref}}) ]

where the importance weight ratio is defined by:

ρi,j(θ)=pθ(yi,j∣yi,<j,cτ,n)pθold(yi,j∣yi,<j,cτ,n)\rho_{i,j}(\theta) = \frac{p_\theta(y_{i,j} \mid y_{i,<j}, \mathbf{c}_{\tau,n})}{p_{\theta_{\text{old}}}(y_{i,j} \mid y_{i,<j}, \mathbf{c}_{\tau,n})}

Here, ∣hi∣|\mathbf{h}_i| denotes the harness token length, ϵclip>0\epsilon_{\text{clip}} > 0 is the clipping boundary, and prefp_{\text{ref}} is the frozen Stage-II repair checkpoint. Unlike standard GRPO objectives that reward candidates solely based on relative within-group ranking, Evo-GDPO explicitly rewards candidates for overtaking historical archive incumbents and achieving greater efficiency under reward parity.

0

1

Updated 2026-10-02

Tags

Prep Sessions

Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Ch.2 Multi-Stage Harness Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Stage III: Evolutionary Group-Decoupled Policy Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor