Formula

Decoupled Reward and Efficiency Formulations in Evo-GDPO

In Evo-GDPO, candidate harnesses are evaluated against an archive incumbent with baseline statistics (br,bℓ,bκ)(b_r, b_\ell, b_\kappa). Rather than combining raw performance and resource metrics into a scalar reward, Evo-GDPO decouples evaluation into three distinct unnormalized channels:

Rirew=ri+λevo[ri−br]+R_i^{\text{rew}} = r_i + \lambda_{\text{evo}} [r_i - b_r]_+ Rilat=I[ri≥br][bℓ−ℓˉi]+R_i^{\text{lat}} = \mathbb{I}[r_i \ge b_r] [b_\ell - \bar{\ell}_i]_+ Ricost=I[ri≥br][bκ−κˉi]+R_i^{\text{cost}} = \mathbb{I}[r_i \ge b_r] [b_\kappa - \bar{\kappa}_i]_+

where λevo≥0\lambda_{\text{evo}} \ge 0 controls the evolutionary bonus and I[⋅]\mathbb{I}[\cdot] is an indicator function. The reward signal RirewR_i^{\text{rew}} serves as the primary objective, whereas efficiency advantages for latency RilatR_i^{\text{lat}} and monetary cost RicostR_i^{\text{cost}} activate exclusively when the candidate preserves or exceeds the incumbent's task reward (ri≥brr_i \ge b_r).

0

1

Updated 2026-10-02

Tags

Prep Sessions

Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Ch.2 Multi-Stage Harness Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Stage III: Evolutionary Group-Decoupled Policy Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor