Decoupled Advantage Normalization in Evo-GDPO
To prevent signals of disparate numerical scales from dominating policy optimization, Evo-GDPO applies two successive normalization stages across the decoupled reward, latency, and cost channels:
-
Within-Group Normalization: For each metric , the raw channel score is normalized across the group of sampled candidates:
-
Batch-Level Aggregation and Normalization: The normalized metric advantages are combined via a weighted sum , where weights are non-negative, sum to 1, and enforce reward dominance (). The aggregated advantage is then stabilized across the training batch: where is a numerical stabilizer.
0
1
Tags
Prep Sessions
Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Ch.2 Multi-Stage Harness Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Stage III: Evolutionary Group-Decoupled Policy Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Related
Decoupled Reward and Efficiency Formulations in Evo-GDPO
Decoupled Advantage Normalization in Evo-GDPO
Evo-GDPO Clipped Surrogate Policy Objective
Evolutionary Group Decoupled Policy Optimization (Evo-GDPO)
Archive Admission Criterion for Candidate Harnesses
Proximal Policy Optimization (PPO)
In Evo-GDPO, raw task performance, latency, and monetary cost are combined into a single scalar reward function.
In the primary objective formulation , what does the parameter control?
Explain the architectural consequence on latency () and cost () channels when a candidate harness underperforms relative to the incumbent reward baseline ().
Match each baseline statistic from the archive incumbent to the metric it represents in Evo-GDPO.
In the Evo-GDPO evaluation formulation, the unnormalized reward signal serves as the ___ objective.
Order the mathematical operations used to compute the latency channel for a candidate harness.
Decoupled Advantage Normalization in Evo-GDPO
Learn After
In Evo-GDPO, which condition must the aggregation weights satisfy to ensure reward dominance when combining normalized metric advantages?
The small positive constant epsilon_num added to the standard deviation denominators serves as a numerical stabilizer.
What primary problem does the two-stage advantage normalization in Evo-GDPO prevent during policy optimization?
Explain the role and mathematical mechanics of the batch-level aggregation and normalization stage in Evo-GDPO.
Match each mathematical symbol from the Evo-GDPO normalization framework to its description.
In Evo-GDPO, within-group normalization is applied across three decoupled channels: reward, latency, and ___.
Place the operations of Decoupled Advantage Normalization in Evo-GDPO in the correct sequential order.
Evaluate which configuration satisfies the required reward dominance condition of Evo-GDPO, and explain why the other configuration fails.
Evo-GDPO Clipped Surrogate Policy Objective