Formula

Decoupled Advantage Normalization in Evo-GDPO

To prevent signals of disparate numerical scales from dominating policy optimization, Evo-GDPO applies two successive normalization stages across the decoupled reward, latency, and cost channels:

  1. Within-Group Normalization: For each metric m∈{rew,lat,cost}m \in \{\text{rew}, \text{lat}, \text{cost}\}, the raw channel score RimR_i^m is normalized across the group of GG sampled candidates: Aim=Rim−mean({Rjm}j=1G)std({Rjm}j=1G)+εnumA_i^m = \frac{R_i^m - \text{mean}(\{R_j^m\}_{j=1}^G)}{\text{std}(\{R_j^m\}_{j=1}^G) + \varepsilon_{\text{num}}}

  2. Batch-Level Aggregation and Normalization: The normalized metric advantages are combined via a weighted sum AiΣ=wrewAirew+wlatAilat+wcostAicostA_i^\Sigma = w_{\text{rew}} A_i^{\text{rew}} + w_{\text{lat}} A_i^{\text{lat}} + w_{\text{cost}} A_i^{\text{cost}}, where weights are non-negative, sum to 1, and enforce reward dominance (wrew>wlat+wcostw_{\text{rew}} > w_{\text{lat}} + w_{\text{cost}}). The aggregated advantage is then stabilized across the training batch: A^iΣ=AiΣ−meanbatch(AΣ)stdbatch(AΣ)+εnum\widehat{A}_i^\Sigma = \frac{A_i^\Sigma - \text{mean}_{\text{batch}}(A^\Sigma)}{\text{std}_{\text{batch}}(A^\Sigma) + \varepsilon_{\text{num}}} where εnum>0\varepsilon_{\text{num}} > 0 is a numerical stabilizer.

0

1

Updated 2026-10-02

Tags

Prep Sessions

Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Ch.2 Multi-Stage Harness Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Stage III: Evolutionary Group-Decoupled Policy Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor