Learn Before
In Evo-GDPO, within-group normalization is applied across three decoupled channels: reward, latency, and ___.
0
1
Tags
Prep Sessions
Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Ch.2 Multi-Stage Harness Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Stage III: Evolutionary Group-Decoupled Policy Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Related
In Evo-GDPO, which condition must the aggregation weights satisfy to ensure reward dominance when combining normalized metric advantages?
The small positive constant epsilon_num added to the standard deviation denominators serves as a numerical stabilizer.
What primary problem does the two-stage advantage normalization in Evo-GDPO prevent during policy optimization?
Explain the role and mathematical mechanics of the batch-level aggregation and normalization stage in Evo-GDPO.
Match each mathematical symbol from the Evo-GDPO normalization framework to its description.
In Evo-GDPO, within-group normalization is applied across three decoupled channels: reward, latency, and ___.
Place the operations of Decoupled Advantage Normalization in Evo-GDPO in the correct sequential order.
Evaluate which configuration satisfies the required reward dominance condition of Evo-GDPO, and explain why the other configuration fails.
Evo-GDPO Clipped Surrogate Policy Objective