Formula

Stage I Joint Customization and Preference Loss

The training objective for Stage I in JIT-Agent couples supervised imitation learning with reference-anchored preference optimization into a combined loss LI(θ)=LIgen(θ)+λprefLIpref(θ)\mathcal{L}_{\text{I}}(\theta) = \mathcal{L}_{\text{I}}^{\text{gen}}(\theta) + \lambda_{\text{pref}} \mathcal{L}_{\text{I}}^{\text{pref}}(\theta) with λpref≥0\lambda_{\text{pref}} \ge 0.

The supervised fine-tuning loss LIgen(θ)\mathcal{L}_{\text{I}}^{\text{gen}}(\theta) maximizes the log-likelihood of validated teacher tokens to teach protocol-valid scaffold generation:

LIgen(θ)=−E(τ,πψ,Cτ,Eτ,hteach)∼DI∑j=1∣hteach∣log⁡pθ(yjteach∣y<jteach,cτ)\mathcal{L}_{\text{I}}^{\text{gen}}(\theta) = - \mathbb{E}_{(\tau, \pi_\psi, C_\tau, \mathcal{E}_\tau, h^{\text{teach}}) \sim \mathcal{D}_{\text{I}}} \sum_{j=1}^{|h^{\text{teach}}|} \log p_\theta(y_j^{\text{teach}} \mid y_{<j}^{\text{teach}}, c_\tau)

The preference loss LIpref(θ)\mathcal{L}_{\text{I}}^{\text{pref}}(\theta) biases candidate generation toward scaffolds that are simultaneously effective and resource-efficient by anchoring against the frozen Stage-I SFT model prefp_{\text{ref}}, weighted by the value margin Δval\Delta_{\text{val}}:

LIpref(θ)=−E(τ,πψ,Cτ,Eτ,h+,h−)∼DIpref[Δval⋅log⁡σ(βpreflog⁡pθ(h+∣cτ)pθ(h−∣cτ)−βpreflog⁡pref(h+∣cτ)pref(h−∣cτ))]\mathcal{L}_{\text{I}}^{\text{pref}}(\theta) = - \mathbb{E}_{(\tau, \pi_\psi, C_\tau, \mathcal{E}_\tau, h^+, h^-) \sim \mathcal{D}_{\text{I}}^{\text{pref}}} \left[ \Delta_{\text{val}} \cdot \log \sigma \left( \beta_{\text{pref}} \log \frac{p_\theta(h^+ \mid c_\tau)}{p_\theta(h^- \mid c_\tau)} - \beta_{\text{pref}} \log \frac{p_{\text{ref}}(h^+ \mid c_\tau)}{p_{\text{ref}}(h^- \mid c_\tau)} \right) \right]

where σ\sigma is the logistic sigmoid, βpref>0\beta_{\text{pref}} > 0 controls preference sharpness, and log⁡p(h∣⋅)\log p(h \mid \cdot) is length-normalized sequence log-likelihood.

0

1

Updated 2026-10-02

Tags

Prep Sessions

Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Ch.2 Multi-Stage Harness Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Stage I: Task-Conditioned Customization and Preference Learning - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor