True/False

In the Stage I preference loss LIpref(θ)\mathcal{L}_{\text{I}}^{\text{pref}}(\theta), the reference model prefp_{\text{ref}} is actively updated alongside the policy parameters θ\theta.

0

1

Updated 2026-10-02

Tags

Prep Sessions

Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Ch.2 Multi-Stage Harness Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Stage I: Task-Conditioned Customization and Preference Learning - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor