Evo-GDPO Clipped Surrogate Policy Objective
The training objective for Evo-GDPO is formalized as a PPO-style clipped surrogate objective driven by the decoupled, batch-normalized aggregated advantage and regularized by token-level KL divergence:
where the importance weight ratio is defined by:
Here, denotes the harness token length, is the clipping boundary, and is the frozen Stage-II repair checkpoint. Unlike standard GRPO objectives that reward candidates solely based on relative within-group ranking, Evo-GDPO explicitly rewards candidates for overtaking historical archive incumbents and achieving greater efficiency under reward parity.
0
1
Tags
Prep Sessions
Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Ch.2 Multi-Stage Harness Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Stage III: Evolutionary Group-Decoupled Policy Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Related
Decoupled Reward and Efficiency Formulations in Evo-GDPO
Decoupled Advantage Normalization in Evo-GDPO
Evo-GDPO Clipped Surrogate Policy Objective
Evolutionary Group Decoupled Policy Optimization (Evo-GDPO)
Archive Admission Criterion for Candidate Harnesses
Proximal Policy Optimization (PPO)
In Evo-GDPO, which condition must the aggregation weights satisfy to ensure reward dominance when combining normalized metric advantages?
The small positive constant epsilon_num added to the standard deviation denominators serves as a numerical stabilizer.
What primary problem does the two-stage advantage normalization in Evo-GDPO prevent during policy optimization?
Explain the role and mathematical mechanics of the batch-level aggregation and normalization stage in Evo-GDPO.
Match each mathematical symbol from the Evo-GDPO normalization framework to its description.
In Evo-GDPO, within-group normalization is applied across three decoupled channels: reward, latency, and ___.
Place the operations of Decoupled Advantage Normalization in Evo-GDPO in the correct sequential order.
Evaluate which configuration satisfies the required reward dominance condition of Evo-GDPO, and explain why the other configuration fails.
Evo-GDPO Clipped Surrogate Policy Objective
Learn After
In the Evo-GDPO clipped surrogate policy objective, what does the reference distribution represent?
In the Evo-GDPO objective, candidate policies are rewarded solely on the basis of relative within-group ranking.
In the Evo-GDPO clipped surrogate loss formula, what specific quantity is denoted by ?
Explain how the importance weight ratio is constructed in the Evo-GDPO clipped surrogate objective, identifying the probability distributions in its numerator and denominator.
Match each mathematical component of the Evo-GDPO objective function with its role or definition.
The Evo-GDPO objective regularizes policy updates using token-level ___ divergence computed against a frozen reference checkpoint.
Order the mathematical operations performed to compute candidate 's token-level clipped surrogate term before averaging across the harness length:
How does Evo-GDPO treat Candidate A relative to Candidate B, and how does this behavior contrast with standard GRPO?